Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2412.05237.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:08:10.041348Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:07.154168Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c46585a1-cbc9-40e3-9432-9fe4c0339963 · inbound
Visual Large Language Models for Generalized and Specialized Applications MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 294
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02c1b66-f08c-445d-a481-2881b210d20a · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b93518cb-dc78-4422-81fb-50abce96ca5d · inbound
Qwen2.5-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d37b096c-0952-4d52-aa29-37be6b437a11 · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0ccbae95-297b-469a-988f-6eacb9334bb8 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7df89b55-a208-4c96-8197-b601b33a9e54 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ef85ad1-efd7-4ce8-b69a-92e49a222d70 · inbound
SmolVLM: Redefining small and efficient multimodal models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 954eea10-c87b-43cf-9ec2-6b1c8ad09be5 · inbound
Emerging Properties in Unified Multimodal Pretraining MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 703e8ac5-584c-48aa-bd45-2f71d209e61b · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74420e64-e9bf-490e-9bad-8518905b5848 · inbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 535af07f-bec3-4081-8631-57ba21dd4d66 · inbound
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f099e45b-584b-4429-90f8-7bfef4ee4b09 · inbound
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc616f9-4037-4362-bdce-c9897d81045c · inbound
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e5b393c-95b4-44a1-8b11-cd619c3751bb · inbound
Multimodal Tabular Reasoning with Privileged Structured Information MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b51c26-c286-46df-9016-544557ae8b24 · inbound
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545bdbd5-21ae-42ef-983b-af63ca588c00 · inbound
VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456f258d-cdda-45e3-a431-c54fb8de5b8a · inbound
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de573936-c557-4175-a5a9-a506d29443bb · inbound
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6c5b47-e72a-42bd-911d-b726c8c2e9cd · inbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4fd369-42ff-48dd-aa16-8a3b2a588772 · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f42d94-ef92-4dc4-9138-292463bcf64b · inbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb502f8-5f62-4b9c-91fe-4b1e13d4cbbe · inbound
A Survey on Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b685855f-c7de-430b-a29e-3eb89d776736 · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262f9d64-69a0-4f8f-bde8-be33d8710e31 · inbound
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cdb59a1-27ed-4549-b01c-685666eba19b · inbound
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 422dbeb9-5e7f-4fe7-90ca-88f2e6ae6291 · inbound
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbf47d8-68f0-498c-9092-dccac2e046ba · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f5eeb8b-5a5a-402f-8a76-d3fcc9128ac7 · inbound
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cd4543e-76bc-4aef-82f1-3d065e399684 · inbound
Meta-CoT: Enhancing Granularity and Generalization in Image Editing MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 76be5cd6-6450-49f2-94c9-332626a6307e · inbound
ZAYA1-VL-8B Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e757b67-7e18-4a2b-886d-e421665c2846 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 17807cb5-4354-418f-bf11-a74de5b5aafc · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d4359dae-d5c6-4551-9dd7-9321babad716 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6fce37bc-1f7d-4864-8b6a-863019af4870 · inbound
Zamba2-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f408383a-4775-42a0-928c-b751d250b4af · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53f7f1b4-06df-4382-85c3-14791267f844 · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a48d31d2-7767-4a25-86a5-01e2722079ac · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.