Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2407.06438.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:27.734866Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T16:50:10.233691Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b0c3328d-b256-43ab-91ea-11c8c42329ab · inbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2dcaa5-bf61-40e2-8bce-da9ee0cf63ce · inbound
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d27094e-ac80-4478-9cc7-737b3866e3d9 · inbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1b00db-8e89-435d-aa31-4194798f1758 · inbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fee6e8-3154-41aa-9617-1ab5e12fcb06 · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d783e8fa-cc1c-4ef8-b4d6-8b90a045d9f8 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a0b0cc-42e7-4957-8077-93ec3a6cb291 · inbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545f009c-c476-415b-8c41-ec954171c74e · inbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · inbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c525bd1-5ec4-4b4f-9005-06ab4929f066 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b02216c0-a23f-4b29-9223-7e859dbb5a5b · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d5c32b-d851-49e6-af21-ed1bce506d36 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.