Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:21:31.276109Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2411.08840.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:21:31.276109Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e4a1a8b6-d1a1-42d3-8707-fc29bd1b28b7 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models More specifically, we dynamically match the optimal aspect ratio from a pre-defined set of aspect ratios
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 001472b2-f566-40b5-b1e3-d5030bd4ffe4 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142eba1e-b3b4-41f4-a772-c423b5157bee · outbound
Multimodal Instruction Tuning with Hybrid State Space Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7160a4-22f0-41d1-b35b-6272f9fff9e0 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5578964-dc53-41e4-a039-6108e00a76bb · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Decoupled Weight Decay Regularization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad92f513-c98b-40e1-af4f-49b289e651c4 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c88423f-f7fb-41a1-9612-d905940e17da · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6716bf1-354e-4cdc-9498-48cb07466df0 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95542b91-fafe-459b-b498-accacdcdc2e2 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d23893-0e01-46ba-9887-341cb2ad38ee · outbound
Multimodal Instruction Tuning with Hybrid State Space Models CogVLM: Visual Expert for Pretrained Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d15deb-c9c8-4ac0-b63d-6ea572fc1d9e · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Jun Xu, Tao Mei, Ting Yao, and Yong Rui
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b9b1e8e-77b6-4105-8b12-dbc809a0b116 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4989d8fa-ab77-4f91-9f8f-634b06f8d027 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b51e270-0819-4985-b7db-2b6c60a43f96 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd98e160-5018-4e77-874c-99b88dc14132 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models OtterHD: A High-Resolution Multi-modality Model
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f66227-56ea-4375-98be-376e58cebb61 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models A diagram is worth a dozen images
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de71c57-d720-4808-a494-35eb29de45a8 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Chartqa: A bench- mark for question answering about charts with visual and logical reasoning
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 272e81ab-a02a-4c4b-a2c1-78a7eb4e10a4 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be1c1a8-d2c4-4693-b890-d6e2937923ed · outbound
Multimodal Instruction Tuning with Hybrid State Space Models CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006412cd-00d2-4ff4-8d96-09bfdd895ee3 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c228e30-e9ff-4550-b0a7-45cf1fc1c0d2 · outbound
Multimodal Instruction Tuning with Hybrid State Space Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.