Pith. sign in

Paper Citation Record · LEDGER

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs

As of 11 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2602.01554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.01554 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T08:09:19.209759Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact26
  • verified fuzzy33
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58ce4b98-3151-403e-8678-b5cb699491c3 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.804597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:353c58eda4c2fdbeaf7e0c1d11c2f86d7f871e443f3c73e982f6a11cdb933219

Observation 8fd7f759-33e7-475c-9192-79f74b909354 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Emerging Properties in Unified Multimodal Pretraining

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.471135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:81a2cfabcb8c5b119217393c174882ceac571eff9920591b1ab46979d628b5d1

Observation a87d0b9f-bee2-4c42-8203-bd23922d0974 · outbound

This paper cites VILA-U: a unified foundation model integrating visual understanding and generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs VILA-U: a unified foundation model integrating visual understanding and generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.802514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:69683646162e41d25f059bc7718d7950be537efb23b54bcba795956e04c4235f

Observation 6b2134b2-ea71-4322-a8bd-a6ba1af188b7 · outbound

This paper cites Harmonizing visual representations for unified multimodal understanding and generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Harmonizing visual representations for unified multimodal understanding and generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.791680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:f27dbcd241e2459ccaa988b5d13716cd555ae9c60e68934bc4ea25ee35633835

Observation a93ba93a-5f62-4594-ac70-4734fdbff175 · outbound

This paper cites Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.442284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:3fbfa06fb41122a9209b54e9cffad761bca677942925601a9ad86873fc1547d1

Observation bce5ebb7-6b17-4854-9c77-0b558575a297 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.429366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:d3aabc992473e3748f35f334baf983306030fd355b2c058c5f4f2b5bc277c71b

Observation 16d3aa9f-6f7b-4885-a065-99e9b5a2a79f · outbound

This paper cites Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.419989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:9f7468ea5088ea63e0cad3f1dfc53838249f766e9bbf86152b0dafdecac4dc76

Observation 21aa7041-6a31-430c-b826-375928f15e2f · outbound

This paper cites OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.393690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:512982d0e70918c4c5a91545136872d5b7d05ee402393bc6590a9537c7a5e204

Observation 34d74abe-0497-4b11-98c8-41dec6369d04 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.460863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:d8ab6b73fb2d9b5649964a7652c95f0b9cb686d7f2ffc73bd26e0009162b6cd8

Observation 668b7147-1e13-4c54-9923-d88565289cde · outbound

This paper cites Show-o2: Improved native unified multimodal models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Show-o2: Improved native unified multimodal models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.787432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:38e15f73aaccafdd200165b9ce38827f7e3d65887805c6c01cf872202195fae7

Observation 6e7991a9-3bd4-4bd2-a89c-3acd107ba0c1 · outbound

This paper cites UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.390395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:11e6882f0bc7852fbc75f229728fdc04c7a52f72cfffe1a194a6a95d62ab53c4

Observation 81f9b42e-26d5-4a11-9a60-2812f2d2863b · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.783435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:6a26bcf9b354082c34914895f982eb0f696ff98e6acf1cbc280758c3f82f79d6

Observation d0261ed9-428e-4761-a745-e34cadb8a499 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.432288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:1d53c21aa70e3db6161a1ba7a230daaf279b67fa50fa2919098c49a44618915e

Observation 2766ebe7-d804-4519-8ceb-57a85d2c874c · outbound

This paper cites Evaluating object hallucination in large vision-language models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Evaluating object hallucination in large vision-language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.793651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:9d6cde2c1d86f9d6a5669cef9651654fea3baf29908b2b4a8e5d1c5c99d817ed

Observation 03658430-ab94-4bfe-a54a-5941d7bbec70 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.464350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:4cfcc71b4fbffbffd6d6793186ca9fe14deda3b442174d98165032dfbf24fee0

Observation 1bdb6ec7-6418-4641-8b43-ca29e325fd60 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.789390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:60f9e1d128350a4faabebdafb423015e0d8abd8df9f765380dbbde6724abb563

Observation f2d72a31-ca80-4701-8aa3-cd41697ea66e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.736202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:ee9f72f78024e83fda398963a78b979b72c2790ff6bbe92d973b88d30196c1ff

Observation 929f172a-908a-4279-b0b6-c6542b092d9a · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.738791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:de14cf516b3687ca8c2c5c2491b8e05db61794060841b722a0da5aa0c6195a01

Observation ee11affc-1452-4348-b2b3-13d6e7e85949 · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.467647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:1b2d3e4a915c0c6dec13f9237cfbfeb0c5d389fb015b04d2c1a7b4ad414a162c

Observation 9c2aa275-6840-49f1-90e7-785eb62e1f75 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.411067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:3eddef530b499500e7f74b18b820a373b1255644b67a9b4bc576c21b474bc1f7

Observation a53038e8-c08c-4504-a691-32bedb2fb814 · outbound

This paper cites The information bottleneck method.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs The information bottleneck method

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.396288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:ddad56452ebb4df9d317c12bcf042644f2e589e46548982fd4cd99169d17b8aa

Observation f1dae431-31f9-43d3-b9d3-86beda9ed022 · outbound

This paper cites Deep learning and the information bottleneck principle.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Deep learning and the information bottleneck principle

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.785420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:9cd929b957edc39778a275a76d7d6a05d7ac614725893cbc2b85c747429c0c9b

Observation 5f270857-db46-4691-b3cd-05e95079b6c2 · outbound

This paper cites Deep variational information bottleneck.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Deep variational information bottleneck

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.797983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:6fdfbefe6b5d9e36193e6a6904d8c9a91386d7a5cdc840d8e59edb4d53611c39

Observation 6d2af541-7330-4959-86cd-3ef4665d954e · outbound

This paper cites Revisiting hilbert-schmidt information bottleneck for adversarial robustness.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Revisiting hilbert-schmidt information bottleneck for adversarial robustness

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.778435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:a3886ee503db9b70aa215069809ccebc1aeb08c40371f3d95b4fe0eaff59c851

Observation dd344693-43cb-4877-8981-26a61d400178 · outbound

This paper cites A survey on multimodal large language models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs A survey on multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.774354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:0b7f1bc97336b0650205cb97f75c63329584906f8f9199a06298c5f756b8341b

Observation aa14fd1f-e248-4680-8992-01f9c3587c10 · outbound

This paper cites A survey of multimodal learning: Methods, applications, and future.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs A survey of multimodal learning: Methods, applications, and future

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.746935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:75f971fb22550980e17229982d6f6316896c603b30f1436bc501804cd8dfe08c

Observation cf934318-b89e-4d22-b114-ca0794ed1192 · outbound

This paper cites Visual instruction tuning.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Visual instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.740783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:878e2b887fc3024147a3cfe1f878f80819aa891937a35852257a291349ab046c

Observation fddb9239-9cc8-4a17-bd6b-46eed9030fc5 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.776534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:dd2cbc96307ae421c69edff9d223a353925c19774dc77c066fd6d6dcd80b5cd8

Observation 8193229d-05ee-4515-b296-442f52f70349 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.772232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:35d5abecf43fbb0ba5946f17850cd524160d0db48a8eb17861232d7981684055

Observation 2ba4ec7f-9de0-4ef6-b083-86933d59ea7b · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.766406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:2f7a28af3726742c7bf93ac4728f24a2b25a899c7d863d49448816c43ca33d66

Observation 927d64dc-afd7-4b8a-ac62-8e4309b7a1b5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Flamingo: a visual language model for few-shot learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.770343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:dee5b8faa4f19d6f0a7a135928b184053b7baf5837634c57a3c83480854184dc

Observation 2fecd58e-8f7b-43bc-9af0-6c7c104f465a · outbound

This paper cites Querying as prompt: Parameter-efficient learning for multimodal language model.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Querying as prompt: Parameter-efficient learning for multimodal language model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.763905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:f02f3615dbb959eb65c3834c87d4a03f3e4cecbed02b522676ad337c3781d9a5

Observation 8da7d960-e4cb-4eaf-8c4e-894884976619 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs High- resolution image synthesis with latent diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.762022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:7c7839336b113fb5d04f77eb094342aad3f308e9da403221f797f481ffc8ad60

Observation 96f3a5a0-a04d-4527-a82f-c758e20b5081 · outbound

This paper cites Scalable diffusion models with transformers.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Scalable diffusion models with transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.768350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:9017368feea0de281ea44db75db78674e9e65671720f9636f0949c37aad1af59

Observation a1cea758-471d-4a47-ab00-4f94eb9e6cc7 · outbound

This paper cites Text-to-image Diffusion Models in Generative AI: A Survey.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Text-to-image Diffusion Models in Generative AI: A Survey

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:10:45.445484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:846c6b0468876a042a9da3ccc5a192766566fd829aabf9592454572f1551828a

Observation a5ac7940-a0d4-41ff-b608-f50d1645afd1 · outbound

This paper cites Diffusion-4k: Ultra- high-resolution image synthesis with latent diffusion models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Diffusion-4k: Ultra- high-resolution image synthesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.780830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:948dbeb63b02e9dbcb28ffebb9cbdcc51dae1162bd28ec476c151b92dc2cba47

Observation bbb82042-0c04-4499-b489-258e7ed117b3 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.405370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:b53dd6ad393ff6274d1e22bd0dec9561f99ab42a13b9db24b870c501aa931311

Observation c74b4f1c-bf74-4b2a-9e7b-fb2557d0ca09 · outbound

This paper cites Editar: Unified conditional generation with autoregressive models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Editar: Unified conditional generation with autoregressive models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.755586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:b1cc397520d7319a3f68387f528369ab7c7b299c288351582ffcc87e5a5baec1

Observation 162561b8-e3bc-4dc0-9713-cf16e2e016ae · outbound

This paper cites Dreamllm: Synergistic multimodal comprehension and creation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Dreamllm: Synergistic multimodal comprehension and creation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.757683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:2c5f5ff71f698c2b4781af53aa987f4f365b102908b9800b85790fa8367c05d4

Observation f8c49b79-879b-4924-8937-811ff85de83f · outbound

This paper cites Making llama SEE and draw with SEED tokenizer.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Making llama SEE and draw with SEED tokenizer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.753465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:f012087e18d36a3bf5738b21d659d10b2636a1489d6f63101ab14458c5bc5590

Observation 0811652f-8afe-46d0-8b30-a277f7896ff6 · outbound

This paper cites Generative multimodal models are in-context learners.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Generative multimodal models are in-context learners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.759650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:8f6c21634088e1ec358cad80258dec635b1d8c33ea68e6203dc78c0ef2c6937f

Observation dd5d5985-839d-4291-93b7-1e67630c33c1 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.438190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:9f5ddebca1acfc04e9c6c19aee7cf48d17c68eab7051f07bbbc576377a22efd7

Observation 71341f95-6b0e-4021-8355-fc1a89ee0e1d · outbound

This paper cites Fast Autoregressive Models for Continuous Latent Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Fast Autoregressive Models for Continuous Latent Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.400026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:048aef28cebf3a828d7505630d1e8843e09067dc508923ad952aff397a5a1c79

Observation c73ac3a2-42b1-45da-becf-909d8e3220b9 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Emu3: Next-Token Prediction is All You Need

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.453837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:5b04f73b7408f686156735ea3a3dd954cf3ab5038588149c929b2f42029ade99

Observation bac26ac1-be08-415c-9dd1-0396a829c3e6 · outbound

This paper cites MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.435592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:a86b9594d183e7a77c70cacad5e13b713d865053bbd45d6971a9ff14c793971a

Observation 9ce7f5c3-b503-4397-8d88-24575dcd997c · outbound

This paper cites Growing visual generative capacity for pre-trained mllms.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Growing visual generative capacity for pre-trained mllms

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.423109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:7ebbf19a989c01d16f07cb7a27ec2f751cd65c2385495c72f77a62df66d422bb

Observation 96348008-d968-42c8-8185-94a74ee37817 · outbound

This paper cites TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:10:45.426729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:9387f0e31c5d84e391e4266cac2be95e4a1f93af495a637f0d617607540d8f9d

Observation ce74419a-cfe0-465d-bdc3-0d5c16ebb065 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.402600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:66558605dd4d1c220e81e8b8c91b3ecf525cf21dfbc5ef5da69675c6125f304f

Observation 9564809d-8d73-4090-86f8-aff930f992f4 · outbound

This paper cites Qwen2.5 Technical Report.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Qwen2.5 Technical Report

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.383518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:1d539cb7d97c4e6a203e7cb4c3cfd263ccea3c317a56feb7acb00a789cc4b4f5

Observation 1b850180-1e66-439a-bdf5-77e147eef537 · outbound

This paper cites Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.387048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:17bcb9153478bd6fd0eef43870d7325e8a9dfeee6aaba617120227b5944df48e

Observation 7647c093-2e30-4c30-9f63-723d791a198c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.457354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:c3b564a156e737a619fb19c265794f6bdc4ea822e25457bf430915ec7a80b775

Observation 8b9f8907-0f25-46dc-a9a2-5c41beab01f6 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Show-o: One single transformer to unify multimodal understanding and generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.799979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:732413206c63af6b8b7b14b5a5029d233ba8d892259bf5b31cb003704d22acae

Observation 4e99df32-0fe6-4ef7-946f-5a06718290ea · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.408310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:5641dd3688e4d1de4cfe1f9e7104474648ec0bb51c699c8a7e26ec8b1bb7bb3d

Observation 6cac06bd-8777-46db-b8de-10d3e71de603 · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.795676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:cc9d751d6d52d4eab39186f42987e1e0170a3b59ba2dd957c76407f5ef1f26fe

Observation 4812055c-79b7-4188-82ca-1604e579d2d9 · outbound

This paper cites On variational bounds of mutual information.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs On variational bounds of mutual information

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.749404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:4955f6f13aa98498b0d179894422ef25bbc63b22b071d232046aad612013c066

Observation a3d8fa67-8548-42b5-a608-00604d9ea937 · outbound

This paper cites Auto-encoding variational bayes.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Auto-encoding variational bayes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.751544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:71cc5634082ddb1bf2cbc11de2af5bcb3f46015ced5e49885fc50fd5c9a6f2c4

Observation 5e83a1f2-e6a0-4416-bfc0-3efcd66b233b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Representation Learning with Contrastive Predictive Coding

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.416560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:b30d456c41bdd4451684b01c46daabe144e03c37be3cd8c9c15fa9ecd91414d1

Observation 301102f7-0ed1-4c7a-9843-725a2aef8112 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Imagenet: A large-scale hierarchical image database

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.745019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:2422fb2a4f9e5f680e47787911a7dc86ba5d7a86592a8824b75a82b53a79b849

Observation 1418c44d-ca66-42ac-90b4-96d8302f9c50 · outbound

This paper cites Densefusion- 1m: Merging vision experts for comprehensive multimodal perception.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Densefusion- 1m: Merging vision experts for comprehensive multimodal perception

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:10:45.742991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:aeb3ca4e75162837695d6e3f1b6ec5b4715abcae826c6d4ad2d485720f3ae70c

Observation 388417b9-e942-4903-b2e5-7420ab00a381 · outbound

This paper cites Ovis-U1 Technical Report.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Ovis-U1 Technical Report

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.413995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:d904c24b9acd4555792158acec415ee8ff8887d0b0f223175af27861863ed44e

Observation 61790730-dd59-45b0-b2ef-d2984ff3aae1 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:10:45.449189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:b8da96de86667cb9283ceac4e574175a3d6e9a3322b837ff2415ad057191011b

Pith citing papers

No inbound Pith citation observations are available.