Pith. sign in

Paper Citation Record · LEDGER

Multimodal Instruction Tuning with Hybrid State Space Models

As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2411.08840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08840 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:21:31.276109Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4a1a8b6-d1a1-42d3-8707-fc29bd1b28b7 · outbound

This paper cites More specifically, we dynamically match the optimal aspect ratio from a pre-defined set of aspect ratios.

Multimodal Instruction Tuning with Hybrid State Space Models More specifically, we dynamically match the optimal aspect ratio from a pre-defined set of aspect ratios

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:21:31.882011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:21:31.276109Z digest=sha256:c4540d673d3b5c2ef3ace941ea8adf1c53e6ba01d697846c4b110a21ab44abbf

Observation 001472b2-f566-40b5-b1e3-d5030bd4ffe4 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Multimodal Instruction Tuning with Hybrid State Space Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.157370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.157370Z digest=sha256:31b0f4cb0a748ea7325902f1691249867a97f59f32d178420b11bd777831b6d0

Observation 142eba1e-b3b4-41f4-a772-c423b5157bee · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Multimodal Instruction Tuning with Hybrid State Space Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.164588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.164588Z digest=sha256:292b458d13e635acfb74274c19ad946f00e3d8f76fbceaa073c8aa59f19ac627

Observation af7160a4-22f0-41d1-b35b-6272f9fff9e0 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Multimodal Instruction Tuning with Hybrid State Space Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.195621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.195621Z digest=sha256:903c0ff12a0e07274ad92636a93270f88d4feccceabdf08d2400c54a03eaecc3

Observation f5578964-dc53-41e4-a039-6108e00a76bb · outbound

This paper cites Decoupled Weight Decay Regularization.

Multimodal Instruction Tuning with Hybrid State Space Models Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.201755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.201755Z digest=sha256:f371a37250492856d465f8375b75bbaa00ff6c6ddafb0b6cd13d4bce5dece76a

Observation ad92f513-c98b-40e1-af4f-49b289e651c4 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multimodal Instruction Tuning with Hybrid State Space Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.208152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.208152Z digest=sha256:00cae78f1ffb9fa56a45b391f1ada74506a7661dc252f4d82181c16ae272ab9f

Observation 2c88423f-f7fb-41a1-9612-d905940e17da · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Multimodal Instruction Tuning with Hybrid State Space Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.214170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.214170Z digest=sha256:7c3cbfe098f977275909e07eb7014e14a12f205c6bca622d8fbb3be5089d33ce

Observation d6716bf1-354e-4cdc-9498-48cb07466df0 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Multimodal Instruction Tuning with Hybrid State Space Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.226420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.226420Z digest=sha256:192211e05986bacf10b6e98a9dc6ad92660ac6b450a6f1e3ad0bb4a76e931981

Observation 95542b91-fafe-459b-b498-accacdcdc2e2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multimodal Instruction Tuning with Hybrid State Space Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.232816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.232816Z digest=sha256:ffc1563a1d9871f18227b9d172ceba5afef5440133549a5f617ce963fd803407

Observation a2d23893-0e01-46ba-9887-341cb2ad38ee · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multimodal Instruction Tuning with Hybrid State Space Models CogVLM: Visual Expert for Pretrained Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.238700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.238700Z digest=sha256:1abbab151240c5b90d6ee001b94e8c96ca5bdea6af532998db9c6e3ef81d5cc8

Observation d0d15deb-c9c8-4ac0-b63d-6ea572fc1d9e · outbound

This paper cites Jun Xu, Tao Mei, Ting Yao, and Yong Rui.

Multimodal Instruction Tuning with Hybrid State Space Models Jun Xu, Tao Mei, Ting Yao, and Yong Rui

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.245629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.245629Z digest=sha256:eabfe31cd2aa9efb1bd12e983853aa32515a36163a4585c0a4cf86995065ca7c

Observation 8b9b1e8e-77b6-4105-8b12-dbc809a0b116 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Multimodal Instruction Tuning with Hybrid State Space Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.260276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.260276Z digest=sha256:b01c73db6c269b2c11efd893b8cad4eea5311dea583afee6649aa5b6e6801f21

Observation 4989d8fa-ab77-4f91-9f8f-634b06f8d027 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

Multimodal Instruction Tuning with Hybrid State Space Models Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:21:31.910068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:21:31.267556Z digest=sha256:a907ce1437417153b416af43899cb46ccf706c9e5dfe44067a79b35a0d61f4f9

Observation 2b51e270-0819-4985-b7db-2b6c60a43f96 · outbound

This paper cites Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg.

Multimodal Instruction Tuning with Hybrid State Space Models Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:21:31.933268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:21:31.252469Z digest=sha256:953044a6fa3b4aad6b6556148e08ea7fc9f44fd361e208303cc709e19392a1c0

Observation fd98e160-5018-4e77-874c-99b88dc14132 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

Multimodal Instruction Tuning with Hybrid State Space Models OtterHD: A High-Resolution Multi-modality Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.183098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.183098Z digest=sha256:cb4e62a4f3154a821bb528ccf0d2f549ecfdf3f1d0436bf6e24623af3f9ca413

Observation e5f66227-56ea-4375-98be-376e58cebb61 · outbound

This paper cites A diagram is worth a dozen images.

Multimodal Instruction Tuning with Hybrid State Space Models A diagram is worth a dozen images

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.177235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.177235Z digest=sha256:24779498b4063be44215d0094344e33dbeba0dc0887ff90eafc7dc3fa01ad5b5

Observation 1de71c57-d720-4808-a494-35eb29de45a8 · outbound

This paper cites Chartqa: A bench- mark for question answering about charts with visual and logical reasoning.

Multimodal Instruction Tuning with Hybrid State Space Models Chartqa: A bench- mark for question answering about charts with visual and logical reasoning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:21:31.952091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:21:31.220432Z digest=sha256:45095203b91501ef1b6a06b89386d1f9b66086b0fd1f790e3285903e143b6464

Observation 272e81ab-a02a-4c4b-a2c1-78a7eb4e10a4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Multimodal Instruction Tuning with Hybrid State Space Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.171280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.171280Z digest=sha256:3d660a715ecfe65271769c28ba9f7c8b03ac2a0c4e7497f425f79b804b4ece5f

Observation 7be1c1a8-d2c4-4693-b890-d6e2937923ed · outbound

This paper cites CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts.

Multimodal Instruction Tuning with Hybrid State Space Models CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.189198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.189198Z digest=sha256:3d8ccf38a2cdd7981312fcb48434a0e7ca8135fe70aec1cb340fb889fd5a1804

Observation 006412cd-00d2-4ff4-8d96-09bfdd895ee3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multimodal Instruction Tuning with Hybrid State Space Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.144628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.144628Z digest=sha256:3539bc1e029795ff8e949389ea9ae1d81ce96e0daf92279a758c27eb9a583a22

Observation 9c228e30-e9ff-4550-b0a7-45cf1fc1c0d2 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Multimodal Instruction Tuning with Hybrid State Space Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T21:21:31.150782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:21:31.150782Z digest=sha256:ef668f09fd79403af75c5ac4d71060a2d57b8c8e2f00ff1410fa4789360aad43

Pith citing papers

No inbound Pith citation observations are available.