Pith. sign in

Paper Citation Record · LEDGER

Unified Multimodal Understanding via Byte-Pair Visual Encoding

As of 7 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 4 inbound Pith citation observations for arXiv:2506.23639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23639 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:57.394029Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T12:10:53.720348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact3
  • verified fuzzy14
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f10ad92-c19c-4b6d-9332-606a508c5f8a · outbound

This paper cites A Survey on Multimodal Large Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:50.412114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:50.412114Z digest=sha256:41ee56a5471cb8aa3da384357bee478ae98f19a6a941d7e1db3513cf5f5c57c4

Observation ac1ecf7c-2e98-4186-b378-21b5978d53cc · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:50.528861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:50.528861Z digest=sha256:22914f4b8de366f690724e32778c3472127e655b42748a022b4dbba7ed6b02c3

Observation 97bfa309-ed67-4035-a364-4fff9c880e7d · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Unified Multimodal Understanding via Byte-Pair Visual Encoding EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:50.645391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:50.645391Z digest=sha256:639f3d20669a3010d64d447e03770bdfa0589983dd89a06fdc402f9a6fc64c07

Observation 94ccf16c-1968-4899-ad09-671fccbaec06 · outbound

This paper cites Multimodal machine learning: A survey and taxonomy.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Multimodal machine learning: A survey and taxonomy

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:50.786380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:50.786380Z digest=sha256:d2f87e3f9ee8fba66470064a3d8359c6f3d7d1cb52e6fe05595497dc0b5557b4

Observation 60152022-7859-45d1-8304-401b38e3e0be · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

Unified Multimodal Understanding via Byte-Pair Visual Encoding How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:50.893822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:50.893822Z digest=sha256:e2d2c64878d345108f7ef985efc5c50dc9204111a51a40956dda124bb5f767fb

Observation 62a8891b-c809-48cf-b8fa-8d0059b3e8d1 · outbound

This paper cites Can MLLMs Perform Text-to-Image In-Context Learning?.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Can MLLMs Perform Text-to-Image In-Context Learning?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:58.277662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:50.990343Z digest=sha256:5a82b80f7cda70fb7edbdddc3bfd0be7b1db037dec1fb5efd1b6715ecf318e8a

Observation 5a1e2ff2-eeb0-438f-a881-2b65ee2bb72e · outbound

This paper cites Vision transformer with quadrangle attention.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Vision transformer with quadrangle attention.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:42:00.700109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:51.077788Z digest=sha256:f58d0aed4b1f9c8c87e2bd2b7d570d08a27ff5444d1c320366b4b8ebbab11da9

Observation dceb408a-7db8-4a2d-81ec-511e00dab3db · outbound

This paper cites Unified language-vision pretraining with dynamic discrete visual tokenization.arXiv preprint arXiv:2309.04669, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified language-vision pretraining with dynamic discrete visual tokenization.arXiv preprint arXiv:2309.04669, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.188244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.188244Z digest=sha256:bae2af653799b03900f4eff71b03a8fa5dff4bb27535155021db903b74a652cf

Observation a2c3ae77-95fe-4fe4-a46a-9825521193aa · outbound

This paper cites VideoOrion: Tokenizing Object Dynamics in Videos.

Unified Multimodal Understanding via Byte-Pair Visual Encoding VideoOrion: Tokenizing Object Dynamics in Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.372359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.372359Z digest=sha256:fa1830c5ae24079448969089b70128bb016f45c08640fc318f3ab7556ae25e81

Observation a3db720f-9570-438f-9fc3-f03e80fbcb8a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Learning transferable visual models from natural language supervision

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.514992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.514992Z digest=sha256:179c265d25918090c29655dc91f5e164108ed886d3edfbca685a032cc2044d72

Observation 40c626f2-5ee3-4f25-95ef-a2428b0320ff · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.615891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.615891Z digest=sha256:041829bd0109b284d344d287ef6f3c238d50ca6223d2f1d3e1187a77e476e605

Observation 14df6c30-c4ee-4b25-8cef-7c9e5cd860f6 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.726124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.726124Z digest=sha256:7e4db07e1d2d0898b00289534e8beebc4fcd8f1e6a438a85dcf9bfd43db33d2c

Observation e7959311-e83a-4d77-ab7c-fac51c3246bd · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.797613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.797613Z digest=sha256:2edf2566719c0403389fffadf746915e8c7c75bfe7085f025ff5007b98dbe0e3

Observation c05a16ac-6e9d-4f4c-971b-e8f615e5a5c7 · outbound

This paper cites UniCode: Learning a Unified Codebook for Multimodal Large Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding UniCode: Learning a Unified Codebook for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.878757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.878757Z digest=sha256:02091e83fead51769d6e68e5cf8921d09482081a600618973073ae3e50591ea6

Observation fb917842-eb1b-4b0d-9024-15256782abc8 · outbound

This paper cites From pixels to tokens: Byte-pair encoding on quantized visual modalities.

Unified Multimodal Understanding via Byte-Pair Visual Encoding From pixels to tokens: Byte-pair encoding on quantized visual modalities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.980688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.980688Z digest=sha256:2c3f97780255ecb65603a9f1d22c738a51be1b2a8f0ca8ad05e60b83dbdf6be0

Observation dbae2431-b06b-437d-ba45-76293f103156 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Neural Machine Translation of Rare Words with Subword Units

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:52.044917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:52.044917Z digest=sha256:687927d82a0b304cc788efd2de109a399dd8304a9569e9765584d132065721d1

Observation 32554849-e934-45eb-9793-572d81adae65 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:52.179859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:52.179859Z digest=sha256:0f72cd9b30c995220b39e20bc25e6d171dcb7cb715be7d43311188f1ddd4bd42

Observation 3aa4c49d-5a16-4c3d-8688-d66f604f5dc2 · outbound

This paper cites Theoretical Analysis of Byte-Pair Encoding.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Theoretical Analysis of Byte-Pair Encoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:52.290519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:52.290519Z digest=sha256:aa996a281b27dbd8d44cb37d04ecddb9400f7f9c0b0e23aa00a85f8de0fefc9f

Observation 904da95a-eb36-46a0-a7e4-908e7b5f86fb · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 2017.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Attention is all you need.Advances in Neural Information Processing Systems, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:52.416594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:52.416594Z digest=sha256:13546cd7017d144e2c36bcb49c4cab80ee8b6911e59d57c7f9c9ea5716ab4518

Observation 378860cf-f482-455b-8674-b479503e0e1e · outbound

This paper cites Sigmoid loss for language image pre-training.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Sigmoid loss for language image pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:52.581194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:52.581194Z digest=sha256:e73fdaa8a3b7b2f65cf43e4da223769c694234c4470ac73b5ed843037d31932e

Observation 908e3bf0-d01f-4182-8f0c-6fe187982334 · outbound

This paper cites Vitae: Vision transformer advanced by exploring intrinsic inductive bias.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Vitae: Vision transformer advanced by exploring intrinsic inductive bias

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:42:00.454518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:52.763886Z digest=sha256:8db4c02ab995b37e9d20e2686b3ee2e807044ba15d96bb0e6959f4596ae7c872

Observation 85e6494a-1958-449a-b4fe-15ba1a6f0e11 · outbound

This paper cites Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond.International Journal of Computer Vision, pages 1–22, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond.International Journal of Computer Vision, pages 1–22, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:42:00.101269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:52.846325Z digest=sha256:08c5953e75f2c6c5deb65f6fc2f9456b80917671c99877fcf4100df7e438e378

Observation 3092ac9c-8f69-41a3-8edc-562323d971d7 · outbound

This paper cites Improved baselines with visual instruction tuning.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Improved baselines with visual instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:52.938876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:52.938876Z digest=sha256:e6ffc8cdf87070148adb53fdc0b874a3aded07432a3f43ec2e5c717da592aafc

Observation 3bb0f35c-9ca9-485e-ad51-fb9a0cc4e820 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Emu: Generative Pretraining in Multimodality

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.036347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.036347Z digest=sha256:caa6db226636ef2e321efb5333da240b1c69cc5dce58aa96b4e045eaa9e1cab4

Observation 4041ed2b-084f-4f77-a239-b461f30b494a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Emu3: Next-Token Prediction is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.151460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.151460Z digest=sha256:cbaf877a8c1ea11db630e9f26dbbce8ae14848f23f815bdf3d959f73175a21f6

Observation 096c485d-06cc-4c21-bcd4-a00f13172303 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.268525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.268525Z digest=sha256:5c451893559cffbfec00ab949d37f30026f14949ec826db23405e85f9ae4b7ed

Observation b3724ed4-3bc6-4ad2-871a-39f8e79c6dce · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Unified Multimodal Understanding via Byte-Pair Visual Encoding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.387232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.387232Z digest=sha256:4bc4bf07e76065d6bdf654cd161065fec0180a0b427d927f24eba85e0c04915e

Observation f78c38b3-34c1-48ad-8341-62f1348a8be0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.485037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.485037Z digest=sha256:88e31687be35359e040395926f59131f5fcb90a981df8259ba96592f9d165369

Observation 727709ad-1e12-44a2-b593-e034d2d5b8ea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.577296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.577296Z digest=sha256:cd0fa4fe613e8089afb1af70acc54f381f2f7477108b8c6bdccb61bc0a8f92be

Observation 0a063002-233b-4342-ab32-2ea48966291b · outbound

This paper cites Qwen2.5-VL Technical Report.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Qwen2.5-VL Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.671687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.671687Z digest=sha256:a2981f874bfeadeea1154774fa2237b204bebf658cd5236745ff32b848b75bbc

Observation d60b045a-5505-4938-a6af-2c0fb6f8b28a · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Unified Multimodal Understanding via Byte-Pair Visual Encoding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.737982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.737982Z digest=sha256:88fb0fc948bb50b012bc6e2478880a73f85231f12519c8f4b4290cfa4c618724

Observation dfa9bd41-838a-417c-8080-dccc295891f5 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Unified Multimodal Understanding via Byte-Pair Visual Encoding LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.815994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.815994Z digest=sha256:99d4fab14172451ac2058830c789551b11a4df84114dadbbefb79a5536c0a4bd

Observation df313483-fe6a-4964-a8f7-6a58ae9e0d70 · outbound

This paper cites From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities.

Unified Multimodal Understanding via Byte-Pair Visual Encoding From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.953513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.953513Z digest=sha256:960584d5dd3944c9307000d896e2064ef2f3ac36f2af6d88aeab49096b295498

Observation e8525320-212f-4275-b23c-269bad6a5898 · outbound

This paper cites A systematic literature review on multimodal machine learning: Applications, challenges, gaps and future directions.Ieee access, 11:14804–14831, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A systematic literature review on multimodal machine learning: Applications, challenges, gaps and future directions.Ieee access, 11:14804–14831, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.945916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:54.056540Z digest=sha256:e38a1ae3376dd98b32d2ccdc96bb56e280084074ebe381d0ee02e4182197a4b7

Observation f975a922-0aaa-4b32-b46c-2f3a787244e7 · outbound

This paper cites Vision language models are blind.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Vision language models are blind

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.165235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.165235Z digest=sha256:180a21b037d3697ca0b9888c4e6790671cd18868c3a5787166499e5f01a57863

Observation dc7a07c4-4dc1-4b6b-971d-e4ca46e75119 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Hallucination of Multimodal Large Language Models: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.265985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.265985Z digest=sha256:a3d05919552c87fe5b4fc1069f5202c9520845c7d9105b06a2f985d14380f907

Observation 490c5c7a-309b-4952-9dd1-f72c7038452c · outbound

This paper cites Visual Hallucinations of Multi-modal Large Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual Hallucinations of Multi-modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.374403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.374403Z digest=sha256:38d802eedb33364606778f6a3f0fd44b0a30d55d328e42c81c3f38b844456352

Observation ad0bdb8b-76c9-4534-8da4-0fe5fbeff5ad · outbound

This paper cites Cognitive Mirage: A Review of Hallucinations in Large Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Cognitive Mirage: A Review of Hallucinations in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.444370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.444370Z digest=sha256:d59f18472bd67036ce05cbe1eaa3c19fdd616a732aced00a1543207aa3fcf276

Observation 7d1c231f-6acf-4df2-a2fd-6eec4c216bde · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Taming transformers for high-resolution image synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.525356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.525356Z digest=sha256:bbca82e534d81ccb2b9a91f2682560f1d0c68c819883006b7cbefea3333caa77

Observation 8cd077cc-ae94-4316-a74c-32b7d553a2a1 · outbound

This paper cites Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.592659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.592659Z digest=sha256:65413f977c1bebedb6ebbc3a7c717edeeab71f0e4b5f3384b8c2ec57935cc590

Observation 5edc9fc9-745c-495d-a8e4-b31378f804c5 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Generating diverse high-fidelity images with vq-vae-2

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.650809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.650809Z digest=sha256:c1d0a1685e3a88cd6b892cb20e11b15fcba55e5d008f27713756bd2f4ecfd97d

Observation 289f6d8b-521d-482b-9a11-6bfb53eb8c6f · outbound

This paper cites Investigating the effectiveness of bpe: The power of shorter sequences.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Investigating the effectiveness of bpe: The power of shorter sequences

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.797657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:54.723510Z digest=sha256:ce80485bee1a6aa0214467ebfc03e79901906232b5f31195ba5c2633c584cf5e

Observation 9136c071-c921-468c-b820-01f716a8a1f9 · outbound

This paper cites Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.806449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.806449Z digest=sha256:f19f9306be83c09ce44a8745433d7bec35e773ae3e63c8928f6dc8ed2126dbc3

Observation 4a4cc56b-1ec5-4cc6-8a9d-a9b0568455fa · outbound

This paper cites Language Models Still Struggle to Zero-shot Reason about Time Series.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Language Models Still Struggle to Zero-shot Reason about Time Series

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.902737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.902737Z digest=sha256:a9193d57fcb576ba91ea7f4d1adb3301751b5eca07ca7fe202745375b4094698

Observation 7d9971e9-a33c-4ab5-a4fc-725df0b3063d · outbound

This paper cites Toward a Theory of Tokenization in LLMs.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Toward a Theory of Tokenization in LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.967902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.967902Z digest=sha256:3b966de0174214075c405c972382b28b6e5830208367313d57fd21d96dfc99fa

Observation 99163494-5a32-4d4a-af9d-d3e0e4e63147 · outbound

This paper cites Pixel-level bpe for auto-regressive image generation.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Pixel-level bpe for auto-regressive image generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.606670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:55.035413Z digest=sha256:27ee22fd1c93b4790e4df3c4c88cecd3da9ca88c520ea3d074f74c12e7c49efd

Observation 65e2c676-d829-4538-a94d-e73944121cd5 · outbound

This paper cites Analyzing The Language of Visual Tokens.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Analyzing The Language of Visual Tokens

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.095533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.095533Z digest=sha256:dc6d06cb3d3634f8fa5ef0a8c9b2405c5aab29d648dc0d78f78bafad853157bb

Observation e3fe4b26-24f5-4bad-a652-228daf10189d · outbound

This paper cites A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:57.746183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:55.164896Z digest=sha256:30c157566a4607aede249af424fd45d2f231475662804346090e281e92625b37

Observation 077de308-52d3-442a-bc02-bbce26764b6a · outbound

This paper cites Exploring Selective Layer Fine-Tuning in Federated Learning.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Exploring Selective Layer Fine-Tuning in Federated Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:57.626943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:55.268049Z digest=sha256:889a8c667e6a6c1aa58663186f3ad0bed3dd9ffc0613aface58eb14da37530d8

Observation 981bc695-6f18-4fbd-8167-9004f152bb67 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.442424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:55.325277Z digest=sha256:05a608169ca308af186195020e54acba560a4315c3a980507a8f6959e5ff222f

Observation 12192f30-2fac-4490-97fa-162e39571921 · outbound

This paper cites A Survey on Multimodal Benchmarks: In the Era of Large AI Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Multimodal Benchmarks: In the Era of Large AI Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.375826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.375826Z digest=sha256:90a94045c508a0a57ef61c2e2d96c1c07cb7bc5b8c2de9fefd5add2939771296

Observation 2bd0f30d-8816-4898-adaf-fc8413a2e385 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Benchmarks of Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.436309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.436309Z digest=sha256:5185ed5f6f05900f44e1d78b82fce3397e64fd139e61a3c67c9bb25f7939fdc7

Observation 7d3a3677-1ddb-407e-abe5-0762e4c7c939 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.516529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.516529Z digest=sha256:b10cbbeb24d1dd0790bc8ac92935b4b85921d49e8d3a042866aefb5e534bbfa5

Observation 5bf461eb-40f4-4ff1-b045-c705725af440 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Vizwiz grand challenge: Answering visual questions from blind people

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.606187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.606187Z digest=sha256:d53da21f21a1c636ed57e84c6660e9d5544c2d261fc2504b9711ae7d27d4007a

Observation 72cb88e8-64a8-4d22-bd2e-47e4db4b67fc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Unified Multimodal Understanding via Byte-Pair Visual Encoding MMBench: Is Your Multi-modal Model an All-around Player?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.671714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.671714Z digest=sha256:6b9680bc6507dcb69846e8e52a3ab952c7079a193e5592064c974f5343336b3a

Observation 81ba791c-6e18-4c35-8f3f-9e6042cd0f98 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.809922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.809922Z digest=sha256:af9473b2798350a1f94b296aa95d2881b8b44eaed0aba7d552e5032fe39f5500

Observation 5a3742cd-c88d-4525-8100-e84213ff1ae3 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.904538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.904538Z digest=sha256:1c7fc5e147732ac04296090ee359fc7a205b5447aa5ca3d73ce325cb4eda30ae

Observation d6c8899b-482b-4dc1-8a79-b03640c421f4 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Evaluating object hallucination in large vision-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.319725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:56.006889Z digest=sha256:c71722a570adf7b625f5aad9898cd8ca3ab63f68e282fe634070d2a47e625921

Observation 38ca8925-6387-4d10-93fb-9a00c8f191f3 · outbound

This paper cites Instructblip: towards general-purpose vision-language models with instruction tuning.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Instructblip: towards general-purpose vision-language models with instruction tuning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.178757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:56.146472Z digest=sha256:09894294b4671abc8b6f3316f6573792854a2d7d129c03cd904059ac29307305

Observation 14b2c4c5-2293-4d61-b0c4-41a22caf16cb · outbound

This paper cites Visual instruction tuning, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual instruction tuning, 2023

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:56.195219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:56.195219Z digest=sha256:fd3f3d1e7584702a1aaddac12e3205733ee3b5294458061471de1fcf212a4c6c

Observation b59c49a0-dc63-46cf-967d-66606af31a34 · outbound

This paper cites mplug-owl: Modularization empowers large language models with multimodality, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding mplug-owl: Modularization empowers large language models with multimodality, 2023

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:59.070947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:56.339136Z digest=sha256:421b091f6782a92deb3a9de3db89b3d46bcc48855e81379c27ee6b8a8382c032

Observation 7497e7c1-2a3d-44d7-a0cf-e234318c7f90 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:56.453016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:56.453016Z digest=sha256:664edaa23e731ec3f973a606de0f552cc1cd3109404f4df0892340d2c0c3ddba

Observation c35a48af-9d8b-4ec5-8b89-de100a1fd173 · outbound

This paper cites Hyperllava: Dynamic visual and language expert tuning for multimodal large language models, 2024.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Hyperllava: Dynamic visual and language expert tuning for multimodal large language models, 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:58.920856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:56.541168Z digest=sha256:ebcc45f09371e1e6da6108b2027d738b2e2e3d5206f3882be09d77ed8bf2c198

Observation 0bdf722b-94a7-4898-a7a9-10dd7bf76d1d · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Unified Multimodal Understanding via Byte-Pair Visual Encoding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:56.615410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:56.615410Z digest=sha256:7393b6380aad84eb86a1a40b741cb6cddd23d032a6de7ebed9052f98cb15585f

Observation 0d5f0ccf-bdad-4abd-aa59-8982c2369309 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Vila: On pre-training for visual language models, 2023

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:56.711512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:56.711512Z digest=sha256:20c778088d5313af4875ede1ee9eb638637b11a27c726e038b8032a7ff876e77

Observation b6543114-bace-461b-889d-350cb30c3cfe · outbound

This paper cites The Llama 3 Herd of Models.

Unified Multimodal Understanding via Byte-Pair Visual Encoding The Llama 3 Herd of Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:56.795379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:56.795379Z digest=sha256:60d12ffd8fc699d68d89cef972a7ccf58851af5e21ace729eda4945864f8a566

Observation b98bea3a-bbb0-48c6-81be-9b3e47a38762 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:58.779225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:56.942368Z digest=sha256:31fe26c3bad6d73765b94ba263ddf48e25c6460f7690673cef569f3d91f73503

Observation e4ee348a-c6c8-4e94-bbd9-24ef3c999c51 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advancesin Neural Information Processing Systems, 35:25278–25294, 2022.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Laion-5b: An open large-scale dataset for training next generation image-text models.Advancesin Neural Information Processing Systems, 35:25278–25294, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:58.633479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:57.024279Z digest=sha256:962d7e09f5e63082a9f69aa8096fe52fe6b9da163348678fd5e6e5e0c8e59811

Observation eb35cce0-617a-4291-8d8a-ba3cbe3e9ca5 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Referitgame: Referring to objects in photographs of natural scenes

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:57.082546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:57.082546Z digest=sha256:5f4e1978e177b793bd0b19be47c7831dceb4318c84928b29a90efed26af9d92a

Observation f070fbc9-7d31-495a-b523-63b483c7784c · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A-okvqa: A benchmark for visual question answering using world knowledge

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:57.138710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:57.138710Z digest=sha256:c58c6a550e7eb67f3fa5c7b45a33dbcc5d7745e5e72a3533278678934103f857

Observation 01018029-1a25-490c-90db-70a372e72504 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Unified Multimodal Understanding via Byte-Pair Visual Encoding LLaVA-OneVision: Easy Visual Task Transfer

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:57.184367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:57.184367Z digest=sha256:3cbed2ce61be3329c4c3d585c3246f11163f8367832db0c9894324cf465c4e04

Observation 4af90841-7a56-4b60-89ce-66678354582c · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:57.241233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:57.241233Z digest=sha256:65cf5c07962f7d1272a3a5cf018baad8b0dc2e90a0576932687783ef42ad6ca6

Observation 4cb5fe32-b9e9-42ac-9d5d-df4e41a384bb · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:57.327402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:57.327402Z digest=sha256:b2690b082af39645b63e1ef791322a3c80d17c3de5592ab399ec3bef1e1171e9

Observation a65df49f-9baf-4d94-89dc-31f0ad9adbc7 · outbound

This paper cites • Reasoning Data (RD): We utilize 504K general QA entries and 343K reasoning-focused entries from the LLaVA-OneVision Dataset [71].

Unified Multimodal Understanding via Byte-Pair Visual Encoding • Reasoning Data (RD): We utilize 504K general QA entries and 343K reasoning-focused entries from the LLaVA-OneVision Dataset [71]

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:58.451432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:57.394029Z digest=sha256:11b2f4a64866db817d972b2ef95d6dc21ffcd480ffb36a1f6b6143ebc76edd47

Pith citing papers

Observation 0e248dc9-df63-4d2e-8dbc-eb43593f295e · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Unified Multimodal Understanding via Byte-Pair Visual Encoding

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.550981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:36c9a7d7e291d380f5e7df13538ed2556ff19e457149cde0333881f9af13ac9f

Observation 6f14d2a7-9d87-4387-94ef-c078c35a3c65 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Unified Multimodal Understanding via Byte-Pair Visual Encoding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.438120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:35df3e50119061bfc7ca892aa90a32a0e5830f3d77250148d876cd8d87735e9d

Observation 33d616b9-24d6-49f7-b3c2-152668849488 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Unified Multimodal Understanding via Byte-Pair Visual Encoding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:b1953afcabc64dc152a3ed134fafe03120e8b1cd75077f586bce44cfbcd633f8

Observation 8cb96130-6798-494e-8c20-cfd5a500a861 · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos Unified Multimodal Understanding via Byte-Pair Visual Encoding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:01:04.671383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:2a0f24619c2833cfead654053d8f429c5e9a567fd2f67a4b29221e069f8f2b78