Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T01:23:32.849326Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 1 inbound Pith citation observation for arXiv:2604.20705.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T01:23:32.849326Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:56:43.645518Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T13:56:46.683829Z
88 of 88 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 260c7929-66a1-46dc-886a-60cae45f37e9 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7bc843e-ed67-49d5-b2b5-ee1778e69b07 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Self-supervised learning from images with a joint-embedding predictive architecture
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a27f76a-3d07-4553-94af-98249770fdaa · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 678f4dde-8fc0-427c-b2fb-f565c4cac325 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 519e7bb4-19da-4c46-995d-2253a66a0ee4 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Beit: Bert pre-training of image transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6de91f32-5026-4cbd-9d09-454f11cf5282 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Lan- guage models are few-shot learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ba7cdb8-2fec-4404-a476-cc43c0cc9c8b · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Deep clustering for unsupervised learning of visual features
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29c08d6f-4c75-4f23-8b65-9442b54d5e6e · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsupervised learn- ing of visual features by contrasting cluster assignments
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 893ee7a6-dac2-4328-a9d6-b79e94a853a9 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Emerg- ing properties in self-supervised vision transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1fe64c31-31d5-4297-8359-d454c9701b3a · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Are we on the right way for evaluating large vision-language models? InNeurIPS
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00797619-2f47-4ff5-bc5a-1123148b7e87 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Generative pre- training from pixels
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd787b00-f372-4301-ac3e-37847a5cfde3 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models A simple framework for contrastive learning of visual representations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4b18c12-05ee-4ddb-96b9-8a3b0ce58f19 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation daa06f95-2bab-40aa-872a-89bedc694d43 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a7e73d1-be31-4e56-ab9b-096a84c2f5be · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation deb9e839-53b7-4a34-8173-beb01eb4a784 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsuper- vised visual representation learning by context prediction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b950c390-7cfb-4074-bba5-af510f3531bc · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models The Llama 3 Herd of Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5840e98-8e44-4e93-be20-49f813b53cb8 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 038b08ac-291e-422f-b352-997195d78f74 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Eva: Exploring the limits of masked visual representa- tion learning at scale
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cfa8c795-34ce-4f2c-8ea6-2b0b928eb6b9 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Un- supervised representation learning by predicting image rota- tions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88020b11-fa72-4ea7-9fe6-3b6b80d346a8 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Bootstrap your own latent: A new approach to self-supervised learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b772da9-4a27-433f-9759-c0559c37e9f7 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f04a502-80b2-4872-8fd8-e9d950ff0a1f · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ceb19fab-9b3d-4fa7-b481-43a4ba4d7ade · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e3c9789-eb81-4034-b0eb-06be2026dfe4 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Momentum contrast for unsupervised visual rep- resentation learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 794cfba8-0e6b-476a-9ee8-2df42d097097 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Masked autoencoders are scalable vision learners
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c412a58-09fd-4cb3-a362-de815ed439c8 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ecf6064-e608-47a3-ab62-4b1c0329d284 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models OpenAI o1 System Card
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f74776ad-b8d7-4ed0-b424-ff7e276642ea · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9540e78-ff6a-44ec-85e5-7c9ad553a4bb · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Lisa: Reasoning segmenta- tion via large language model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5479c03e-29b8-4f31-b514-cb40dfc63a56 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da42ec44-bc0f-4b05-bf6f-be5f1de66c33 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55580ff7-2933-4045-b829-1c30cf8c13e6 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61c047b1-c908-43ea-b229-df68a73628a5 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Correlational image modeling for self-supervised visual pre-training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75226c71-ce97-427c-87db-a0944021055f · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 45b8ad45-8804-4acc-b3b5-d38344be5b35 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Microsoft coco: Common objects in context
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 450d6001-2a2b-410b-bb2a-dd023c5bc7b9 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual spatial reasoning.TACL
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6b3e8c3-86e2-4ae9-b711-646368e5ae17 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual instruction tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab6c62c3-027b-4bfe-ad40-ed40e9a94902 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? InECCV
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27d1c04d-607c-4cbc-b8a3-76406c449774 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Spatial-ssrl: Enhancing spatial understanding via self-supervised reinforcement learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2fc0114f-1199-4283-b7cf-ccb76b2ae1c5 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual- rft: Visual reinforcement fine-tuning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 809ac577-0d4a-4546-b8e2-90440b037115 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27bbc79b-fc4e-4dad-a127-632a143ef0a3 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d468757c-1a7d-4194-8c31-c2238b417063 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e60766ee-691e-48d1-8319-6abe0e486a1d · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.https://ai
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e061e1d5-45bb-409c-becc-56d03b8844e8 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsupervised learning of visual representations by solving jigsaw puzzles
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f23922af-7a14-4dbe-92d3-277a2d355292 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models GPT-4 Technical Report
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 13f89b78-6918-46b9-a6d3-6560fb050141 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DINOv2: Learning Robust Visual Features without Supervision
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a6b55d36-e588-4736-8ad7-c97b7654a8af · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Training lan- guage models to follow instructions with human feedback
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2714f8db-d069-49b9-ad7d-35547320de7c · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Context encoders: Feature learning by inpainting
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71d9472f-15b4-4786-aa7a-76748b494902 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b26737de-543d-466c-b3d2-b6aaef3be2ee · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Proximal Policy Optimization Algorithms
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b73a474-f45a-45f0-8bca-06235fbfd290 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 80ccf399-67a6-40e1-a917-5b01964db5d2 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DINOv3
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9504a04-894b-4e5a-983c-b2da6a2847d1 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 081620ee-690c-4fc0-86ae-b0827f3d1f50 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Gemma 3 Technical Report
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ab89138-c7f7-427d-bbe9-9b69ab18d67b · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d035e965-e854-4323-89c1-d9a3953a5dc2 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Winoground: Probing vision and language models for visio- linguistic compositionality
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6253173f-d227-4eab-bff6-e5711facdd51 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8519b1fe-bf08-4257-a92d-1d8c83d7f77d · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cb32b16-38b8-4bcf-bf2f-0811bdd26fe1 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70e5b200-4fd3-4d98-9252-9bc780772809 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mea- suring multimodal mathematical reasoning with math-vision dataset
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aaf30619-55bd-4342-a66e-dad0355c4977 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a247ef6-a82d-46aa-b669-1071141cac14 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e35c6d6-165c-4e4b-8bbd-2d5be9181d8b · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Vicrit: A verifiable rein- forcement learning proxy task for visual perception in vlms
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bbacae6-4a7b-4f3b-aaeb-2cb579d8f937 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1681915f-b3bf-406c-afdd-25718f8f9862 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bcb35224-6a9e-4f1d-bb75-7b1284dc1b73 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c7eda69-9da2-41ff-8dac-829d76e372ea · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models V?: Guided visual search as a core mechanism in multimodal llms
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 305dcce8-101f-404f-b91b-67a4ccdb4fc1 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual jigsaw post-training improves mllms
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93ddcea5-acf8-4ffb-8237-018a3f76e2a3 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiMo-VL Technical Report
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3015ae6-3e4b-498f-b5e4-51e6e6d2875a · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsupervised object-level representation learning from scene images
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d01f57a8-f546-4faf-8cae-b5d82cb33a44 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Delving into inter-image invariance for unsupervised visual representations.IJCV
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9775d78e-cf60-4eda-bf4f-b5f9f2cfe3fb · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Masked frequency modeling for self-supervised visual pre-training
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 98d290cd-b7d1-4857-a35b-5c5e73eca5f4 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Depth any- thing v2
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af98848a-06ea-4dda-8b6c-663bf307169b · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models How to evaluate the generalization of detection? a bench- mark for comprehensive open-vocabulary detection
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37c41524-cf77-4589-b7b7-e8f28b166e66 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 013d3648-2590-407f-9943-9e8bc8b7be78 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82fd2cc8-b1da-4349-97f8-e1b19fe59164 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7897cae0-cbe9-414b-a511-49f6176688a4 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 785a44ac-6f1c-4d8a-8882-d8bc0f501496 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Sigmoid loss for language image pre-training
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66c10490-5f20-48f6-a458-0ac81f389d47 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Online deep clustering for unsupervised representation learning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 18f68011-3757-4d0c-9051-27a9782c586f · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 174fe4a7-f5dd-4cbe-b0d1-fcb98e2163d9 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenar- ios that are difficult for humans? InICLR
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43e8db37-78f0-47b6-8ad9-cfb67458cd7c · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7adc2e5-d8fd-4fb6-a0af-f7b9fd059301 · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models ibot: Image bert pre-training with online tokenizer
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8811d8eb-aa12-4f9c-971b-d2fd7cab62af · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4efbcf1-98a0-4bfa-9c6e-6c4dd460d0ba · outbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd97243f-5c04-4656-9560-48fe1ac9104e · inbound
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.