Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:57.394029Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 4 inbound Pith citation observations for arXiv:2506.23639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:57.394029Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T12:10:53.720348Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1f10ad92-c19c-4b6d-9332-606a508c5f8a · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Multimodal Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1ecf7c-2e98-4186-b378-21b5978d53cc · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bfa309-ed67-4035-a364-4fff9c880e7d · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ccf16c-1968-4899-ad09-671fccbaec06 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Multimodal machine learning: A survey and taxonomy
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60152022-7859-45d1-8304-401b38e3e0be · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a8891b-c809-48cf-b8fa-8d0059b3e8d1 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Can MLLMs Perform Text-to-Image In-Context Learning?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a1e2ff2-eeb0-438f-a881-2b65ee2bb72e · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Vision transformer with quadrangle attention.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dceb408a-7db8-4a2d-81ec-511e00dab3db · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified language-vision pretraining with dynamic discrete visual tokenization.arXiv preprint arXiv:2309.04669, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c3ae77-95fe-4fe4-a46a-9825521193aa · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding VideoOrion: Tokenizing Object Dynamics in Videos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3db720f-9570-438f-9fc3-f03e80fbcb8a · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Learning transferable visual models from natural language supervision
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c626f2-5ee3-4f25-95ef-a2428b0320ff · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14df6c30-c4ee-4b25-8cef-7c9e5cd860f6 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7959311-e83a-4d77-ab7c-fac51c3246bd · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05a16ac-6e9d-4f4c-971b-e8f615e5a5c7 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding UniCode: Learning a Unified Codebook for Multimodal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb917842-eb1b-4b0d-9024-15256782abc8 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding From pixels to tokens: Byte-pair encoding on quantized visual modalities
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbae2431-b06b-437d-ba45-76293f103156 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Neural Machine Translation of Rare Words with Subword Units
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32554849-e934-45eb-9793-572d81adae65 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aa4c49d-5a16-4c3d-8688-d66f604f5dc2 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Theoretical Analysis of Byte-Pair Encoding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904da95a-eb36-46a0-a7e4-908e7b5f86fb · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Attention is all you need.Advances in Neural Information Processing Systems, 2017
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378860cf-f482-455b-8674-b479503e0e1e · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Sigmoid loss for language image pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908e3bf0-d01f-4182-8f0c-6fe187982334 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85e6494a-1958-449a-b4fe-15ba1a6f0e11 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond.International Journal of Computer Vision, pages 1–22, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3092ac9c-8f69-41a3-8edc-562323d971d7 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Improved baselines with visual instruction tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb0f35c-9ca9-485e-ad51-fb9a0cc4e820 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Emu: Generative Pretraining in Multimodality
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4041ed2b-084f-4f77-a239-b461f30b494a · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Emu3: Next-Token Prediction is All You Need
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096c485d-06cc-4c21-bcd4-a00f13172303 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Deepseek-vl: Towards real-world vision-language understanding, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3724ed4-3bc6-4ad2-871a-39f8e79c6dce · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78c38b3-34c1-48ad-8341-62f1348a8be0 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727709ad-1e12-44a2-b593-e034d2d5b8ea · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a063002-233b-4342-ab32-2ea48966291b · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Qwen2.5-VL Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d60b045a-5505-4938-a6af-2c0fb6f8b28a · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa9bd41-838a-417c-8080-dccc295891f5 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df313483-fe6a-4964-a8f7-6a58ae9e0d70 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8525320-212f-4275-b23c-269bad6a5898 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding A systematic literature review on multimodal machine learning: Applications, challenges, gaps and future directions.Ieee access, 11:14804–14831, 2023
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f975a922-0aaa-4b32-b46c-2f3a787244e7 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Vision language models are blind
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc7a07c4-4dc1-4b6b-971d-e4ca46e75119 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Hallucination of Multimodal Large Language Models: A Survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490c5c7a-309b-4952-9dd1-f72c7038452c · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual Hallucinations of Multi-modal Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad0bdb8b-76c9-4534-8da4-0fe5fbeff5ad · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Cognitive Mirage: A Review of Hallucinations in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1c231f-6acf-4df2-a2fd-6eec4c216bde · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Taming transformers for high-resolution image synthesis
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd077cc-ae94-4316-a74c-32b7d553a2a1 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5edc9fc9-745c-495d-a8e4-b31378f804c5 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Generating diverse high-fidelity images with vq-vae-2
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 289f6d8b-521d-482b-9a11-6bfb53eb8c6f · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Investigating the effectiveness of bpe: The power of shorter sequences
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9136c071-c921-468c-b820-01f716a8a1f9 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4cc56b-1ec5-4cc6-8a9d-a9b0568455fa · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Language Models Still Struggle to Zero-shot Reason about Time Series
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d9971e9-a33c-4ab5-a4fc-725df0b3063d · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Toward a Theory of Tokenization in LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99163494-5a32-4d4a-af9d-d3e0e4e63147 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Pixel-level bpe for auto-regressive image generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65e2c676-d829-4538-a94d-e73944121cd5 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Analyzing The Language of Visual Tokens
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3fe4b26-24f5-4bad-a652-228daf10189d · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 077de308-52d3-442a-bc02-bbce26764b6a · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Exploring Selective Layer Fine-Tuning in Federated Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 981bc695-6f18-4fbd-8167-9004f152bb67 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12192f30-2fac-4490-97fa-162e39571921 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Multimodal Benchmarks: In the Era of Large AI Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd0f30d-8816-4898-adaf-fc8413a2e385 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Benchmarks of Multimodal Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3a3677-1ddb-407e-abe5-0762e4c7c939 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf461eb-40f4-4ff1-b045-c705725af440 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Vizwiz grand challenge: Answering visual questions from blind people
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72cb88e8-64a8-4d22-bd2e-47e4db4b67fc · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding MMBench: Is Your Multi-modal Model an All-around Player?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ba791c-6e18-4c35-8f3f-9e6042cd0f98 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a3742cd-c88d-4525-8100-e84213ff1ae3 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c8899b-482b-4dc1-8a79-b03640c421f4 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Evaluating object hallucination in large vision-language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38ca8925-6387-4d10-93fb-9a00c8f191f3 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Instructblip: towards general-purpose vision-language models with instruction tuning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14b2c4c5-2293-4d61-b0c4-41a22caf16cb · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual instruction tuning, 2023
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59c49a0-dc63-46cf-967d-66606af31a34 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding mplug-owl: Modularization empowers large language models with multimodality, 2023
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7497e7c1-2a3d-44d7-a0cf-e234318c7f90 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35a48af-9d8b-4ec5-8b89-de100a1fd173 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Hyperllava: Dynamic visual and language expert tuning for multimodal large language models, 2024
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bdf722b-94a7-4898-a7a9-10dd7bf76d1d · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d5f0ccf-bdad-4abd-aa59-8982c2369309 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Vila: On pre-training for visual language models, 2023
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6543114-bace-461b-889d-350cb30c3cfe · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding The Llama 3 Herd of Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98bea3a-bbb0-48c6-81be-9b3e47a38762 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4ee348a-c6c8-4e94-bbd9-24ef3c999c51 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Laion-5b: An open large-scale dataset for training next generation image-text models.Advancesin Neural Information Processing Systems, 35:25278–25294, 2022
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb35cce0-617a-4291-8d8a-ba3cbe3e9ca5 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Referitgame: Referring to objects in photographs of natural scenes
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f070fbc9-7d31-495a-b523-63b483c7784c · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding A-okvqa: A benchmark for visual question answering using world knowledge
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01018029-1a25-490c-90db-70a372e72504 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding LLaVA-OneVision: Easy Visual Task Transfer
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af90841-7a56-4b60-89ce-66678354582c · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb5fe32-b9e9-42ac-9d5d-df4e41a384bb · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65df49f-9baf-4d94-89dc-31f0ad9adbc7 · outbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding • Reasoning Data (RD): We utilize 504K general QA entries and 343K reasoning-focused entries from the LLaVA-OneVision Dataset [71]
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e248dc9-df63-4d2e-8dbc-eb43593f295e · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Unified Multimodal Understanding via Byte-Pair Visual Encoding
Reference 218
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f14d2a7-9d87-4387-94ef-c078c35a3c65 · inbound
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Unified Multimodal Understanding via Byte-Pair Visual Encoding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33d616b9-24d6-49f7-b3c2-152668849488 · inbound
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Unified Multimodal Understanding via Byte-Pair Visual Encoding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb96130-6798-494e-8c20-cfd5a500a861 · inbound
Being-H0.7: A Latent World-Action Model from Egocentric Videos Unified Multimodal Understanding via Byte-Pair Visual Encoding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.