Pith. sign in

Paper Citation Record · LEDGER

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

As of 12 August 2026, this Paper Citation Record lists 100 of 281 outbound references and 2 inbound Pith citation observations for arXiv:2606.13289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.13289 v1

Coverage vector

measured 100 of 281 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:01:07.362430Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:56.912092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:55:28.362396Z

Reference resolution

100 of 281 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch35

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38bdd4f4-4ae6-4cbc-8a8e-6a8b1096aaa4 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.195656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:56c9c745fb580ab32e65b28351e151371a6f05d004a7c51e7ebbb278bda2e39a

Observation a99eaeac-f037-4288-b968-bc0c8fbd28f7 · outbound

This paper cites International conference on machine learning , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers International conference on machine learning , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:a7cd9b75d1267b292d15b57703e85b24ed3bc9b0a0ee7f825a14890fd7ad183a

Observation 02e5cf2a-28bb-488d-ad08-6f83a9478409 · outbound

This paper cites 2023 , url=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2023 , url=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c3f87b113ee4cf25efb5d79ffb09a4b2c0e2867e4c36713146d6a38988dec483

Observation cdf36153-0d10-429d-b56b-7db745fd1a9e · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:0036206f7612fea9aca3a2bd2aca8fe039f5a4b7df0bd3e487e8d6662b64b431

Observation b1914e40-d6c6-407d-9df5-c0e2b7c05c9a · outbound

This paper cites OpenAI technical report , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers OpenAI technical report , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:334427fc737c04e97a9c7330ba705f6247aec4a7d242010b24dad06f5c336dbf

Observation 08509cb8-d874-47df-9501-1200f6c3ecf7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.084172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:3e109635777496dc95653ce5cf23360596b50f54d9aa1a47251c393dbe867633

Observation 7d0ad57d-079d-4584-8e35-429289c7a6d9 · outbound

This paper cites Data Filtering Networks.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Data Filtering Networks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:32.070663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c315d27a29136761cdf0b8f905d5176842cd5529386f612718ff5804e481390e

Observation 53ff37cc-e9a1-4912-9182-ffa657609327 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:420cc079b9209aa46c1385da740bee20b7e303e48ab0994d25ddb352a09201dd

Observation dc11bbb2-f9a0-4fad-b741-701085296d1e · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers BEiT: BERT Pre-Training of Image Transformers

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.065519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:121a9958a1d6406a0a46dd2b167fe195c02dc14254708f58e9f31a9dfd36cec7

Observation ff37c413-f814-4513-942d-bfb1945f2d36 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:a569f6893ae3226116d2c821a7904deac59ccca9e9dc08594508caa690aefdbc

Observation 78317259-bb70-4afc-9c69-fdc911f8532c · outbound

This paper cites From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.055745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:fc03e3a77be6442b5edba4aa506187633e841abf9a4dd77b7b52331f28dc7066

Observation 0d7bb7a9-fab0-4672-9fac-46779c9773a6 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:97c5ce15751280f53dc730c31933edeed4d765d7908c189b1d4525586825f834

Observation edc35e0a-9445-4d58-bf6f-9bc4420c5c4b · outbound

This paper cites Benchmarking Detection Transfer Learning with Vision Transformers.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Benchmarking Detection Transfer Learning with Vision Transformers

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.128646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d5551e205330b280812d3cef3771abb6c58a07c1eb861a32e259b1e50acf710f

Observation 2eebba51-a766-4f17-a24b-dc3f1061cae3 · outbound

This paper cites Proceedings of the IEEE international conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE international conference on computer vision , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:6dc32f68d5b95f26b1368896dffb931ab7213a1fdac19cf7d32a13b809ec2ffb

Observation 381f1b43-cee1-4214-9fbd-71cc8c8626bd · outbound

This paper cites ArXiv , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers ArXiv , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:17cb57a6338ac8c94dcbaf63a3569989f5c82d954a5744f65c910e719a79a76e

Observation e653307c-698c-489a-973d-bb1ded75a413 · outbound

This paper cites ArXiv , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers ArXiv , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:04cac024317b6192570fb9db4f0fbd4b34940af764943a8fcf88d34c51bf02d1

Observation b1efcd99-a458-4e1a-9a66-5ad02414ec2d · outbound

This paper cites 2024 , eprint=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2024 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:656aa86ae792d5697fbc562afc4fe0b8b5cc91ebbc67c5b898781fcba099225b

Observation 41549d01-da07-44bf-8fbf-32f8c137c637 · outbound

This paper cites Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.238647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:03ef2ffd6f0ebb729fffc4cc0891448ad420d3d3edf48136e4cfa1b13b3479e2

Observation 03c4caf3-515c-4b02-b175-1eaa7394f102 · outbound

This paper cites Boosting latent diffusion models via disentangled representation alignment.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Boosting latent diffusion models via disentangled representation alignment

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:29.597324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:0bcebd378d4d20d5d59f98e9bd57163e5f23a6cd25f71aa7aed0e06575ccddd2

Observation 9d3abfb3-0d72-45f6-a516-78a74ab5dfc3 · outbound

This paper cites Towards scalable pre-training of visual tokenizers for generation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Towards scalable pre-training of visual tokenizers for generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.907664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f9caa9652d0a2f00afad72b64bdd79be2367618e93ace8f1d0733ac34b630e74

Observation 3636ed34-81fc-46cd-be6e-e88d4a862b57 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Wan: Open and Advanced Large-Scale Video Generative Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.873847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:30de7edb3930382a6d5875cfe6ff21406b48c7d3ecfe43c0797f715ecd174ae2

Observation f7d6ba4f-720c-45ab-9dfc-d85a1f4922e2 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Diffusion Transformers with Representation Autoencoders

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:29.602800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:8bffca43904740240d11c6ad8659d25e31cf8eb5cd723bff756d66ba2b02ead1

Observation 582b5745-87fd-47e3-9de4-c8010932e8a3 · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:29.616556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:20f0408accd250d80671abe10579728f9dc0c61ff43cd4a198428b3c67bdc00c

Observation 38ef4a00-688e-44c8-8d41-719bbf950efe · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:e557c0097140d9c37f6ffe85ecd4b48e87b585d156a32dc0f53bea7e9af4d880

Observation 8dab6d55-943d-4737-8d4b-44524abb0d83 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:29.647041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:25f9ff3eb628f365480856bf6d0de34acb1e267665544ad880fc1d15b8ef5a36

Observation e115c19a-7ec6-487c-a166-d64469dd806d · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:fb6b66a48ab087ed4f2652f7a914bac367c746bf378c545218c126d027ed7b07

Observation 4abdca51-3c1d-4f56-bf19-faf3386e173c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in Neural Information Processing Systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:ffec5b3d9ccbc80ade43ad5ca7906153708b88da0dd3fb1b13668775498ced10

Observation 2338cc16-d921-461a-a725-387847b30ac5 · outbound

This paper cites ArXiv , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers ArXiv , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:2290727c8fc0b277a08bad5da8343f5b9d2aed61045e971e47d18d42adcb201c

Observation be118554-1f50-4082-89ae-df5a631b0461 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers HunyuanImage 3.0 Technical Report

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:29.579017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:578be074d3d1857bf9d7da356904907e5f7e4d9d858a2193804c7435601aa6da

Observation b7f1b6ff-2dcc-4560-aacb-430297665062 · outbound

This paper cites Latent diffusion model without variational autoencoder.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Latent diffusion model without variational autoencoder

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:29.587652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:cd7e654f90b3fb403e6d26bf885eff97926d25e3f445a66453ecfc7febe43915

Observation d0b9e9da-8cd3-41c8-bef1-7dc9fa1b14e2 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Vector-quantized Image Modeling with Improved VQGAN

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:29.663135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:844fee085ad3830f31299907a2e27c2a6cdd78772f9f666eb1f3afa42df4301e

Observation c5f8e4c3-3d5f-4a0d-a217-7affa712167b · outbound

This paper cites Diffusion Autoencoders are Scalable Image Tokenizers.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Diffusion Autoencoders are Scalable Image Tokenizers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:29.657972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:478dfcfd2b18332f2b1cd5174d27a4547f1e5ebb7a094465e4bbc32239f50cf0

Observation 73b856cd-2489-4a98-a6b8-e50d9e6ec436 · outbound

This paper cites generation: Taming optimization dilemma in latent diffusion models , author=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers generation: Taming optimization dilemma in latent diffusion models , author=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:043dde488cf7e4ff903f702b37e5256a52cc93969b7225a4cd7b6c0d205e5ad7

Observation 294fda3b-4030-4564-84ef-2f2d40ecfee9 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:431f65510ccd894d25f5994e960720bfc237d8fe4c66f627f5b84b3c61e7b22b

Observation fb04f536-e5c6-4bcc-8909-919f7e7233aa · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:9e444aa997ddea7a46ac82bb87db93581e19c073097286951531601b9b135264

Observation b47c3c07-ba25-4c7d-a6cb-35c54c83dfce · outbound

This paper cites arXiv preprint arXiv:2507.08441 , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers arXiv preprint arXiv:2507.08441 , year=

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.917237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d8c2a0dfccd847e63029289f7b03003b342fdc8820b7f09a381494605f94ec7b

Observation 23721df6-51e7-48d1-9b31-dcc526261249 · outbound

This paper cites Advances in neural information processing systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in neural information processing systems , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:15578758b690621b623be2f3b115e7595bf3cf17127a1bcf97a6a1ea73b8d27a

Observation 1929781b-9882-4d01-8de5-f2999afddaa6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in Neural Information Processing Systems , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:b3ce05ca96ebd0c10a88913529fde624cfefdf69a913a1dc51f656ff461a0ae2

Observation 7787fef2-8047-4ad3-a8ac-e50abd4f1246 · outbound

This paper cites Latent denoising makes good visual tokenizers.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Latent denoising makes good visual tokenizers

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.905282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:b169f151e17535bf04e1db7237ca6d00b09a2761b87f339c1336056bf959b158

Observation 49bd37e6-f4be-4269-b367-bccfadf8462f · outbound

This paper cites Advances in neural information processing systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in neural information processing systems , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:3acc2f0ea892152b8ce685e34efb2285050e6d1c83c26ec90d09fa9ba4ec41d9

Observation f71baa52-ab77-49c2-b5b6-9e760ef08680 · outbound

This paper cites Foundations and Trends.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Foundations and Trends

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c982599e79086c963475ec3f22cbe0e9ad538e829b4dcc29e54b0df54864d3d0

Observation 1111e54b-6dcb-417e-b0fd-071e82cd40d2 · outbound

This paper cites Image Processing Algorithms and Techniques II , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Image Processing Algorithms and Techniques II , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:8edcba2964d4d97fb9b581aa5640907bd45b27bad7bfd00ca9fc28683ced1323

Observation 54209dcf-d4c5-420a-9e2a-0b3aa6947666 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in Neural Information Processing Systems , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:fc5a89fb1373d64f52fc7cdb5e387070f4ed62178e6fae4707a75c2e1880d7b5

Observation bd2a13e7-4617-4f89-b924-0d3b765139af · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.889202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:12aea83c5331499ca6fe3defa1ec8f4d3fd551661789abcc9caf6a834d2e9fdc

Observation e9d36f68-0d8f-4812-9bec-57ab91a2194f · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.909903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:431538a8669e319ed5b98fd49e9ed99019636ae7fb8ad89522c5c0d2d20d97f0

Observation 613719cd-5b2a-4053-9b77-3a7047568675 · outbound

This paper cites International Conference on Learning Representations , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers International Conference on Learning Representations , year=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:069bbdabd1a49e898c54e5f9cc87a3d95d2022854451061826c9e1724f2ab919

Observation ecddbb8c-3993-45c6-a945-3b9b2ab7f84d · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:84b808f72158bfe42d45fb388950fa103565a71d9c4242f217029403ae97839a

Observation e3b60c85-bd9d-43fc-9357-5d022a1a207b · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.819341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:a4cae14cf9aba8d7edc5041baff444950e7d2ef584a06fca357a8863a2bfc8cd

Observation c41b4ef4-93ed-4183-a671-cb9380c5377e · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Forty-second International Conference on Machine Learning , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f51cb814db9c9ebbde27947dcfa504e479b1e17fa10ff4128d163330a6ac6268

Observation a93c7eb2-26e8-4d8a-9c5a-d0af8fbf165e · outbound

This paper cites 2022 , eprint=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2022 , eprint=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:71dce23959eb5c8b98dd8e5eace1e6ba3cf7555fd9746610d01874301912587e

Observation c453ecf6-ccf9-48e1-a982-2321c425de44 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:8501253bccccc227da0869b87aec9a5a990c7f5ec9ce27616755a918a4405cd1

Observation 092adf5e-e897-4641-8437-ffafa9b149cf · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.816961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:468bb4738eae1eb7ec86cf82fc044d2b4ecd9c0f8fade4ae46d3b72714ecb31e

Observation e6214d90-8861-4a4c-a0f4-2c1b5bfc5f4c · outbound

This paper cites 2025 , eprint=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2025 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:190a47985a195d287964b3b40a3ba77a772670434efdff28918db2114d4d699d

Observation dd51b75a-c007-4e53-b58b-dd6166d1fa01 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.883801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:e5b7e42d6f834e0bc65705eb745bbf0cbce4835b7d6c39a58aacb9b1abc55aa9

Observation 68408783-f0e0-41fd-9149-ca83de540ac1 · outbound

This paper cites Unified language-vision pretraining in llm with dynamic discrete visual tokenization.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Unified language-vision pretraining in llm with dynamic discrete visual tokenization

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.855777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f5a89791ec3e2065fc74e582b0a600b7dea52ef45ba478c780c76d57aaa7aec9

Observation 9254e3b3-1c2e-4e2b-a23d-dc6f3da66391 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:5db03a4b0b3b63ce202bb5a697729248caedce78c8225b0cc96394d251995126

Observation d1ce572f-4dd1-41f7-8b27-b0e4f573064b · outbound

This paper cites arXiv preprint , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers arXiv preprint , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:3dbceaad993ecbe03d3dd9a5422149a7e42e0b9d91986e50a98ae6dadfaa7a66

Observation 7805d954-f801-4294-a566-f745f9568dd5 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Making LLaMA SEE and Draw with SEED Tokenizer

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.919623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:594a50d9bfd239bae458abd7b8c9519e7a7cf25de62748535bfb1480dc796a90

Observation 0c6aa083-e3d5-41ec-a214-a2276619f066 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in Neural Information Processing Systems , volume=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:6047ffc445516c9cdca529ffd8d235abc368a44dfe1a1758b4a2c269d3dce4a9

Observation 0a6b70da-000d-4b30-934e-013547a0b354 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Emu3: Next-Token Prediction is All You Need

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.912359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:2efa669d687bd25ef05f29754344cbb9f874f4a18aea3490d2966fd75d112db8

Observation 65458b22-33de-4bb8-850e-92fe0ca81b9a · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d89bc8ba925a7e92e3b1dd8b9420f6186c40508f32ac1fd28bbcbf3d099e7868

Observation 977fd9ec-6dd2-431b-a9f4-f3480148eb34 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF International Conference on Computer Vision , year=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:1907abeee5e0a62ddb8e070f25cbec5c70b206694efc7d726969c0f7d40c3283

Observation 3383bfe0-8243-4391-98a0-1699edfd775d · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.860881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:feabb66423fcdeac8fb0d09248058f76e1a87f1c0854c365216d22a6fb68d461

Observation 0f0ae40e-923c-4008-b8ac-c0f24f7e5690 · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.879242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:8940f3f5594127c122ca312fb6703e024f36837ec64b935a1664faaed40f45db

Observation f02cbc2a-3ecd-4a24-9410-4721b4f3ccfc · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:fd6b842825f87bd1041bf4fc62aeeaefd8b3f9151d10a3c958db6c07a5f0ccec

Observation 08882222-8ee0-42b4-ae82-e8e2c6eb3999 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d811ea27b3cf17e89775d2cbfe71ed749cb68fa8ddc202f1c693e587a8f19a32

Observation 8e32b757-65dd-4f02-a9ef-236b15d8ab85 · outbound

This paper cites 2023 , eprint=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2023 , eprint=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:2112f4d2a8c274ab20b6b4dddcaa5c18e01362117ae2dd4de3b32eb1fe4fd71c

Observation 9b7811ee-fe9d-4474-bfaa-5edb250b0eef · outbound

This paper cites 2025 , eprint=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2025 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:2b27b720d90399d17d5d060a200208eac3aff3c0512b3a767d35864ed9ab8fc5

Observation 6d93c88c-7d14-4a5d-ae65-d42304fbf072 · outbound

This paper cites QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.819010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:6ab8dd9b04de72da0ce6be51a019d92d9980acfcc163f6fdaab674b0d1669af7

Observation 56505dc3-693d-40eb-bc29-692615926173 · outbound

This paper cites OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.896947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:91ded45dd6de7306256b085ca723c6b0fcb7e61f64cac9c3624276cd6f0480ff

Observation 0c738ec6-7764-4a1d-a5fd-22cfe2530c27 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.909508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:578a15aee062b419be3cda8f1aca5bd3b4e383b2ab0e1c2c91775f77d0dd218e

Observation 0cebaca9-115c-4a46-9df5-a92158d2c38a · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.886403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:35c4bc9aa5c438eb6f34bcc25923aa59f5587c7c5bd4a6ac39f92a4ab5f06170

Observation f6597786-6122-437f-9948-ce371a6ddebf · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Unilip: Adapting clip for unified multimodal understanding, generation and editing

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.190086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f474d805e2c6da5e3bea29d5c914e3d608cb5d6a891599dca864a7293bc25134

Observation 0677f3a0-72b9-48e3-9c23-e0e8fa80aea9 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:42036c0fa4e371bad7cf2f3f07f4c4de5345f875694b3aca4732af633b39d1a7

Observation 750863bb-62c3-4cdf-815e-4f3281272cf1 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.222192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:fb5f13d7dbe862d6f85cbc65c895481077a16a4553ca76440a4aa18eaf7fa01c

Observation 802d7e2d-d8dd-44e7-b3de-2ddda65cc6c1 · outbound

This paper cites XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:32.058282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:78b040bb3f3a5cb223e042e48ef4ee68836ce67ab1df43950de5c951bf692197

Observation 7eb61c66-8c0d-41de-8f2e-6b6208abac1b · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Finite Scalar Quantization: VQ-VAE Made Simple

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.053215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:599cdb269de66c0b90cd0d00c11ac77a2ba74c565884e48d5e4c57feb2730f37

Observation 2f200e47-713e-4aa5-beb2-522225c5c174 · outbound

This paper cites Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:32.060656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:506221df622430f1011c121457647455525095513bcb63f3119ea6d07823c3ae

Observation 2702eb36-e7e9-45bf-a32f-d60ad9f63b47 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF International Conference on Computer Vision , year=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:097d62e4177f517d58ae1fb7735b3b4320d02b01ee2c348bad525b9892c20b5f

Observation 4cb9248f-9406-4295-ad51-3e7438e6479f · outbound

This paper cites Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.068307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:ce176b878b912a472e0600e9f1744273bcf2a5493a36b286d6c02d3a1cca0462

Observation 17636725-4e5a-4026-b180-1e1a1c2987e4 · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Forty-second International Conference on Machine Learning , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:b01985bbfda38f6af9ba183ac2f55b4dd2a3d0c67e30f50e7a5fd0a1295867fb

Observation b2f6432a-7602-4fae-ae8c-d99b66e7788e · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.043849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:8f7aaf898d81ebf02339ee56c445e429ea886584068f2bb293857c83aa8d6c6c

Observation 1be63069-928f-4b15-8b98-f36a20dfec22 · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:3300234fe740844f66397ba1accadf544fc34524fd5f5d8b1c310f8f23165c74

Observation 84b5e0b9-3509-4e9d-bca9-9d2a67a885eb · outbound

This paper cites 2025 , url=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers 2025 , url=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:dc5159e4900b483d26cc13c0d6748d90649e9c9f0b610866b6235f44110749b7

Observation 7221f5a7-2264-4f00-8acf-956dbc3cc119 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:a7a3e6ec287fd5909cc6c2dc142fe67b65a4b564d937819cf16416930872964c

Observation d0032e54-cc03-4315-9e48-79806020c9e8 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:5f8b1424b3cc3a0f0a5f2a8800abdcbe1fb2e091aab1157ac2fbf0605177e5d8

Observation 13c1b387-1eb9-41e9-b899-4dab3c353bb0 · outbound

This paper cites European conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers European conference on computer vision , pages=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:67d9232c0cc3c223f26d3f5353e141baa59bd7521580149d630174e2f6fe8123

Observation c50e2deb-5af4-4a25-a40d-015970efef42 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Evaluating Object Hallucination in Large Vision-Language Models

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.046198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:1da043234abbad5872b07b988241f321c8a48f048085520276b1538c49c340aa

Observation cd070a70-ef08-4dfa-924b-ca50fadc5394 · outbound

This paper cites nature , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers nature , volume=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:855c606108749985c7fbc17d86eb8e70457270d69015b479db607f7afd552ff8

Observation 2bb8a214-b480-438d-8ad8-f116101d8372 · outbound

This paper cites Proceedings of the European conference on computer vision (ECCV) , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the European conference on computer vision (ECCV) , pages=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d248f69232a9ee1f589fb8bc91a959ac225ec2305a04cf56b9253084ecacf774

Observation afc4c253-6112-4d28-911f-b7793737ffa0 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c68b6221299f25515cc60d9fac18111f66d048524025c46f4d24211b553ffba5

Observation fc4ef71f-64a8-4405-8b4f-590c9b026413 · outbound

This paper cites European conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers European conference on computer vision , pages=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:041638474b24b6357115bdaa306fe662e7ac335c90d0a6ecbb5463fecefac6c8

Observation f1e95262-2d95-41aa-a8c3-8d0ace513f77 · outbound

This paper cites International conference on machine learning , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers International conference on machine learning , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d0c1dde8836966a1fb377a345aaa6d9247486685b9e4df19240bcacbf3af76ad

Observation e683a672-62fb-4d26-b850-ea2426fbc961 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.048574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:2111fa3897831f30ab71d6bf60788d920d0b132629b0e4a6defcf93bfe01367c

Observation 64630893-3eba-400a-ba49-bb6c866a70ea · outbound

This paper cites Internvideo-next: Towards general video foundation models without video-text supervision.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Internvideo-next: Towards general video foundation models without video-text supervision

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.051049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:c435a21adb5259afa0f4a328289a006aa70508f7098da5d2fd36100a78174760

Observation 36fac3a4-3f89-4464-afef-b3b4c1f6c9bc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers DINOv2: Learning Robust Visual Features without Supervision

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.922207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:accd11ae893d4c9f745d58869188a051f2b228ac478964dda0758ac5f237b756

Observation 24e8eaff-199f-4922-a142-f9ee555bd31b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Advances in Neural Information Processing Systems , volume=

Reference 103

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f9d235608c93c6f0a2a43892d4ce7bc936eada225dc01fb5b8165944c14d924f

Observation fd201d84-672c-4b9e-bd3f-8edb0fb4ad87 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 104

Resolution
unresolved
no resolver link, observed 2026-06-27T07:01:07.362430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:e25653e8da07fae1b948848e71397fccb6aec6c31ec7c77332d5ff27bbc12373

Observation 6948834e-a774-4b49-8d4d-7ada40ba6c8b · outbound

This paper cites DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

Reference 105

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.073325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:0d6b20b073be7320ffd4a9c9162451cff2ac314823bb6bbd187b6c66abd6b2e4

Observation 470607fd-e8a9-4ccc-bfa1-82f4ba8a0c71 · outbound

This paper cites arXiv preprint arXiv:2503.06764 (2025) 4, 7, 9, 10, 1.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers arXiv preprint arXiv:2503.06764 (2025) 4, 7, 9, 10, 1

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:32.098762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:bcb6a0f907a9a9a0b8a0bbede9091377e99cddaa948134ee5db3f09eb927e2d6

Pith citing papers

Observation 4a166ade-3560-4c34-9537-15b3524dff4e · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Reference 154

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:55:28.368251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T11:55:28.210340Z digest=sha256:1f9eb05caa3ef0141cc4e5f274eb33821ea3249d8bd23061bc47057bb8c28ca5

Observation a2edc933-347d-4858-8181-67aadbfbbd76 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.912092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.912092Z digest=sha256:cbc871ebb3802031233300435c36ea9661b2f64b1f542e8eb9b175b9b2aa72d0