Pith. sign in

Paper Citation Record · LEDGER

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

As of 6 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 1 inbound Pith citation observation for arXiv:2604.20705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20705 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T01:23:32.849326Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:56:43.645518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T13:56:46.683829Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact33
  • verified fuzzy53
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 260c7929-66a1-46dc-886a-60cae45f37e9 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.520874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:f7165c7e8ecb1aa309f9d2fc55ccb65817daa54009696048e9f7e364a70de33a

Observation c7bc843e-ed67-49d5-b2b5-ee1778e69b07 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.496360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:983d6622ab9bdb8547d47c782bd9124193babb1a53dcf49e96b7d6f336aa8a76

Observation 5a27f76a-3d07-4553-94af-98249770fdaa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.779407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:3b43643de5e2e4b25ce104feac9ab59a2b4172513c637067e4df1e0f71b8e313

Observation 678f4dde-8fc0-427c-b2fb-f565c4cac325 · outbound

This paper cites Qwen2.5-VL Technical Report.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.759358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:ed30fe06209c589471376d1892aff6a06ac5daf0155b88e505506f5292ff2f3f

Observation 519e7bb4-19da-4c46-995d-2253a66a0ee4 · outbound

This paper cites Beit: Bert pre-training of image transformers.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Beit: Bert pre-training of image transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.543294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:5b173a7479b58782b3412744e260ffb21e7dcf14d0437589a296c50c16ee0511

Observation 6de91f32-5026-4cbd-9d09-454f11cf5282 · outbound

This paper cites Lan- guage models are few-shot learners.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Lan- guage models are few-shot learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.500216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:0e4b487324e4adba012b0cbf1e5704d80c2a0a9eafefbce530fd3c856e947539

Observation 1ba7cdb8-2fec-4404-a476-cc43c0cc9c8b · outbound

This paper cites Deep clustering for unsupervised learning of visual features.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Deep clustering for unsupervised learning of visual features

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.486793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:203ab175b0b7d5306dab1c0b22ccb4feadda2e69c53004be6be9a36e75719f00

Observation 29c08d6f-4c75-4f23-8b65-9442b54d5e6e · outbound

This paper cites Unsupervised learn- ing of visual features by contrasting cluster assignments.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsupervised learn- ing of visual features by contrasting cluster assignments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.481372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:3fd40f40bbfbd66b49b767e3b82608ea49e7283d62fcb59061b4014cb377a94a

Observation 893ee7a6-dac2-4328-a9d6-b79e94a853a9 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Emerg- ing properties in self-supervised vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.490749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:b55412f73867dc49186614c31ec8014a6383e3616a4885a9bda7b2ee2a8bf9b4

Observation 1fe64c31-31d5-4297-8359-d454c9701b3a · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InNeurIPS.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Are we on the right way for evaluating large vision-language models? InNeurIPS

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.507615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:a3472ee54178137de08f3f0bc2a68d22e2403da79b426aa8b338775c55cff8ad

Observation 00797619-2f47-4ff5-bc5a-1123148b7e87 · outbound

This paper cites Generative pre- training from pixels.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Generative pre- training from pixels

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.373032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:eee270c72620ddc11b49c1ea2a9d809ca5848fef3ea6a7ede2bce916f8a64f11

Observation dd787b00-f372-4301-ac3e-37847a5cfde3 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models A simple framework for contrastive learning of visual representations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.377078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:fc851c45e2e4f43884e06355ea0e565c4c4aca67a3c3deb0e8aecb38c8dc9346

Observation f4b18c12-05ee-4ddb-96b9-8a3b0ce58f19 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.405150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:a05fb33eba71b5291eef9b53268a0f04f02a623ab6f4c49fcde38cca578002ce

Observation daa06f95-2bab-40aa-872a-89bedc694d43 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.398759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:a3e620b8798b68ad9ac5697f0b3e00af1d31d55e6bd6e84433a1fef51b06f13f

Observation 0a7e73d1-be31-4e56-ab9b-096a84c2f5be · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.409457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:f913c73c6abbb4a0e44b4a96330fc7ef147c3f36898e79808e93904a4096dd5c

Observation deb9e839-53b7-4a34-8173-beb01eb4a784 · outbound

This paper cites Unsuper- vised visual representation learning by context prediction.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsuper- vised visual representation learning by context prediction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.469225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:c9f0c669bdcb6b632d05b7ce062518bf9e11be0ee6b38e57a3a0483e76b03dd1

Observation b950c390-7cfb-4074-bba5-af510f3531bc · outbound

This paper cites The Llama 3 Herd of Models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models The Llama 3 Herd of Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.870137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:3fb8c0b2394bbfa6e0b867abd561745e5594c90fd84eaa7bf02b21ed2f9ed53c

Observation b5840e98-8e44-4e93-be20-49f813b53cb8 · outbound

This paper cites Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.473309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:22bb497f2ee6fb531b8d9b0ef526e177117905965ce9c551b738f65811decdf2

Observation 038b08ac-291e-422f-b352-997195d78f74 · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.512661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:afc4b3b8fbcf2ff503c5fb2c7b5456553695d390391b50e071b81c78ad835a3f

Observation cfa8c795-34ce-4f2c-8ea6-2b0b928eb6b9 · outbound

This paper cites Un- supervised representation learning by predicting image rota- tions.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Un- supervised representation learning by predicting image rota- tions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.516466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:a8ebb625a116ba98404521ddaa8a0d91d973126b26a3ca26132f00cc91cd73b3

Observation 88020b11-fa72-4ea7-9fe6-3b6b80d346a8 · outbound

This paper cites Bootstrap your own latent: A new approach to self-supervised learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Bootstrap your own latent: A new approach to self-supervised learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.539508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:2e9cba4c329f4ad0271eef4b24acb9d1af16ab9f64dcc583177bed9f1a0c2fc9

Observation 3b772da9-4a27-433f-9759-c0559c37e9f7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.659468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:cc19db9dbc42d938d63dc20222ed93a17a0b3cb3a38529884e639def38e09476

Observation 2f04a502-80b2-4872-8fd8-e9d950ff0a1f · outbound

This paper cites SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:02:57.484204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:4b1e6bd05dcb035540b5d9c60edb0b6ef792c9559c098afa0e8e013d747be270

Observation ceb19fab-9b3d-4fa7-b481-43a4ba4d7ade · outbound

This paper cites Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.460941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:639a9d474771552b3e4e0a7ec921f4f35fee4a4b0504a674a586de5076dca81d

Observation 3e3c9789-eb81-4034-b0eb-06be2026dfe4 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Momentum contrast for unsupervised visual rep- resentation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.457155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:4d03c15193ed2d12f35aeb84a00b163f29da84b8111438328d35af3cf558f07a

Observation 794cfba8-0e6b-476a-9ee8-2df42d097097 · outbound

This paper cites Masked autoencoders are scalable vision learners.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Masked autoencoders are scalable vision learners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.391003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:ece42ca3cd8e7433d4cca90c13ef014f5932b90e8c4fe013d6cf770910b66dbe

Observation 5c412a58-09fd-4cb3-a362-de815ed439c8 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.865648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:56b57d859da23a26744e43bff8a0d44e0624d8fe17d1fe45814d450eacad4517

Observation 0ecf6064-e608-47a3-ab62-4b1c0329d284 · outbound

This paper cites OpenAI o1 System Card.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models OpenAI o1 System Card

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.792449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:076a12212aa0bdca7b3f291fcaa9caf633cf0000b1bfcec468cc291481126416

Observation f74776ad-b8d7-4ed0-b424-ff7e276642ea · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.581149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:3ebc79d1a5e4c3a43916c65089fac99e3a4960073fa6f376e4d5caa6df30c1c7

Observation c9540e78-ff6a-44ec-85e5-7c9ad553a4bb · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Lisa: Reasoning segmenta- tion via large language model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.313330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:83af9e792eb3e8db57f98d0e886461320eba73f39fc85d07b2393dbf03d1dee6

Observation 5479c03e-29b8-4f31-b514-cb40dfc63a56 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.726255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:9461bed475298c827f5f7fdf5646843d53f5647876b05fee84e3842e5d0502f6

Observation da42ec44-bc0f-4b05-bf6f-be5f1de66c33 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:48.053680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:a33054bc9d0618da5d7a4ca4e6f386231b7574e5f63467822817c3a9cb8ed82f

Observation 55580ff7-2933-4045-b829-1c30cf8c13e6 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.317375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:0bb528a5c6079c571e0416e6cf335c751a5ca9ef9c30246ee24b6cd0aafbcbb1

Observation 61c047b1-c908-43ea-b229-df68a73628a5 · outbound

This paper cites Correlational image modeling for self-supervised visual pre-training.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Correlational image modeling for self-supervised visual pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.328458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:de66339d48760c6dfe841760c1206f9dcfd6748c6cc1c2cf48fadbb16028f608

Observation 75226c71-ce97-427c-87db-a0944021055f · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.748348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:67ead9461c2c99e16e340c9dacef9397337e6a5abb54e935e4cd08c394f3d6e5

Observation 45b8ad45-8804-4acc-b3b5-d38344be5b35 · outbound

This paper cites Microsoft coco: Common objects in context.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Microsoft coco: Common objects in context

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.362888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:993e56322b6699a9498b9c3ab84310fd8446c282819119fb10b192dc161cd814

Observation 450d6001-2a2b-410b-bb2a-dd023c5bc7b9 · outbound

This paper cites Visual spatial reasoning.TACL.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual spatial reasoning.TACL

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.324720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:5034144c3560b0dd6e4b4d3c30846a34bfa03d6754c75ec651c71a77adbe56eb

Observation d6b3e8c3-86e2-4ae9-b711-646368e5ae17 · outbound

This paper cites Visual instruction tuning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual instruction tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.335615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:57c89ddfe7c249f09608f48d63910ec52797e37f29ed9d2f80bb74a3820bbe96

Observation ab6c62c3-027b-4bfe-ad40-ed40e9a94902 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? InECCV

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.305499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:4e1bb95b4d22fe1a02d551a818c9ea4a257d5c73e3e659acd206bd31995dbac3

Observation 27d1c04d-607c-4cbc-b8a3-76406c449774 · outbound

This paper cites Spatial-ssrl: Enhancing spatial understanding via self-supervised reinforcement learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Spatial-ssrl: Enhancing spatial understanding via self-supervised reinforcement learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.828062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:0ef8ec2f5e02976d78fef5ca71d32903f52563ffbcd09d2fe19bffe0609ab8cb

Observation 2fc0114f-1199-4283-b7cf-ccb76b2ae1c5 · outbound

This paper cites Visual- rft: Visual reinforcement fine-tuning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual- rft: Visual reinforcement fine-tuning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.301849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:7ef015fd2f7d229df6d81e332c72ff3ec9dfd6f180bf0ae076f7bee0903e8918

Observation 809ac577-0d4a-4546-b8e2-90440b037115 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.896552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:520e2e84d07458b8a482fbdc775a4e14948e5348d927a6b92429e7a53dffa702

Observation 27bbc79b-fc4e-4dad-a127-632a143ef0a3 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.298257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:87cef6ad0f87857862862aff28c6b5b284b1dd85ecb75ed2ca4986bbb3aa3e5e

Observation d468757c-1a7d-4194-8c31-c2238b417063 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.597995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:b20ec66436bb03823917f8d8f306b68f3e8048507e48672580baf32f4354fbff

Observation e60766ee-691e-48d1-8319-6abe0e486a1d · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.https://ai.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.https://ai

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.290848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:037b0892cb9702e0e82f357b87ab5bff64fbc27c994ea437c2665b0271186d48

Observation e061e1d5-45bb-409c-becc-56d03b8844e8 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.294569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:8626610ba702ebe0898b89c8d7414d21580e10bfef44bf4c3dfa6dd12f4efebe

Observation f23922af-7a14-4dbe-92d3-277a2d355292 · outbound

This paper cites GPT-4 Technical Report.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models GPT-4 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.810006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:2652841555cd1dbebd493a77d2e344ed8481c1e9f99bc34ed217e86c0190190f

Observation 13f89b78-6918-46b9-a6d3-6560fb050141 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.839114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:9f910ea99bb3f4ab49ab031c3fbd4252b4dd4840ca19d6db12dda1cd82b0f569

Observation a6b55d36-e588-4736-8ad7-c97b7654a8af · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Training lan- guage models to follow instructions with human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.309561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:e7e1fdeba697d894bf0ac66542a0db072fb5b0e789ffc36a89c9c9a571f4c1cc

Observation 2714f8db-d069-49b9-ad7d-35547320de7c · outbound

This paper cites Context encoders: Feature learning by inpainting.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Context encoders: Feature learning by inpainting

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.320920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:d885f7ea37ce39f2b9516d87bc821968612782317714804de369b6b59de6b43c

Observation 71d9472f-15b4-4786-aa7a-76748b494902 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.465153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:8432d8d69b934fdc1e29a9ad2223d35911530f80e1d0d88ca6cb61d7d91870db

Observation b26737de-543d-466c-b3d2-b6aaef3be2ee · outbound

This paper cites Proximal Policy Optimization Algorithms.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Proximal Policy Optimization Algorithms

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.667118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:610d57e5fcedb54ab40a3377a9c4e45e42fa22b20c6414462f5c0bd439689d01

Observation 8b73a474-f45a-45f0-8bca-06235fbfd290 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.851864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:741dbad8ac45982c81d293a91a98eafdf3cddc74966c46d9db951cedc0eb3774

Observation 80ccf399-67a6-40e1-a917-5b01964db5d2 · outbound

This paper cites DINOv3.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DINOv3

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:36:08.572435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:60a24895c6f2b647d4175ff7135a5d58f884d268f85ae48bf9dfc60e8e730ff8

Observation b9504a04-894b-4e5a-983c-b2da6a2847d1 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.576979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:0cff35d37343f64e14e56a20843712aaaf35469e7138ff77177136648b2d9eba

Observation 081620ee-690c-4fc0-86ae-b0827f3d1f50 · outbound

This paper cites Gemma 3 Technical Report.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Gemma 3 Technical Report

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.844575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:9773313d53ecb6e6a2e230fdaea53d4de95ec2333f1e0c3677fa8eaad68d431c

Observation 2ab89138-c7f7-427d-bbe9-9b69ab18d67b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.822073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:e44d3cba56c33c70515e045d5c6c9d38184630e9aa3c39a6c5d59228e524cf78

Observation d035e965-e854-4323-89c1-d9a3953a5dc2 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.477546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:4d1bcd8e56ebc4845c577b8b7a6034425f3b8899a7791ef9ae2caccf740b216e

Observation 6253173f-d227-4eab-bff6-e5711facdd51 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.524656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:a776912c34d9da8e1e636d340562c310227ac77307a265bb7d4109ba0c238e98

Observation 8519b1fe-bf08-4257-a92d-1d8c83d7f77d · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.814689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:bcbfef4bccec81bf46ecbacdc2749609beec9f773c79a0baa576d453596153c8

Observation 4cb32b16-38b8-4bcf-bf2f-0811bdd26fe1 · outbound

This paper cites Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.535693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:cd55230620a81c9a0b01bb56e00f9f5b42af48a45026a738f3e6373bc243d45f

Observation 70e5b200-4fd3-4d98-9252-9bc780772809 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.453066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:0dd7e69eddde1cde172628dc380e543b83834601b3f2e92c67d4bb98fc29c652

Observation aaf30619-55bd-4342-a66e-dad0355c4977 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.445129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:6b40227eac6607e88650782f3a9e9029cd5821ff602a3e9c0d6f800f8416cb6c

Observation 9a247ef6-a82d-46aa-b669-1071141cac14 · outbound

This paper cites LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.805123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:3da2d6cfc4d79c1a44a74ad31f909b45415421fffbb9bcc4e1f797cd1b7f3182

Observation 9e35c6d6-165c-4e4b-8bbd-2d5be9181d8b · outbound

This paper cites Vicrit: A verifiable rein- forcement learning proxy task for visual perception in vlms.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Vicrit: A verifiable rein- forcement learning proxy task for visual perception in vlms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.531901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:65e0375c66fdba50378fd6ed2b8431c5ba04ebdc3f63b12370c83a3686150790

Observation 4bbacae6-4a7b-4f3b-aaeb-2cb579d8f937 · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.856150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:d3d29d2e75eb7d43ec055d6e2442e93f3b7c34a5d4bf5f89404e5e7fcc353a32

Observation 1681915f-b3bf-406c-afdd-25718f8f9862 · outbound

This paper cites Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.652216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:156b111af3fe266d47c21d681cdf3291810dbe7ac2172ee63bcea940b3cc713a

Observation bcb35224-6a9e-4f1d-bb75-7b1284dc1b73 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:27391e915ef5929c13315322b36ceba04ff5147a87178b21ab228c3728ffeb83

Observation 6c7eda69-9da2-41ff-8dac-829d76e372ea · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models V?: Guided visual search as a core mechanism in multimodal llms

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.528152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:463b0a87be5cc3ea188c47eb94442c0bf25e4b962445656441f025aa9a1f2fca

Observation 305dcce8-101f-404f-b91b-67a4ccdb4fc1 · outbound

This paper cites Visual jigsaw post-training improves mllms.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Visual jigsaw post-training improves mllms

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.680333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:e556ab23eb11102c27ece22398b6c282eb86564fa161615879e5d7c9bc3b17cb

Observation 93ddcea5-acf8-4ffb-8237-018a3f76e2a3 · outbound

This paper cites MiMo-VL Technical Report.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiMo-VL Technical Report

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.615360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:6cb99ecf78ddfa948c98a017b93f6b0651bcf561bdae17e07007e037575eb25a

Observation a3015ae6-3e4b-498f-b5e4-51e6e6d2875a · outbound

This paper cites Unsupervised object-level representation learning from scene images.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Unsupervised object-level representation learning from scene images

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.394749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:699f0122da4c1aed741d588fb93b610bb96f7dbbd2bce81068b2addc5ab40e78

Observation d01f57a8-f546-4faf-8cae-b5d82cb33a44 · outbound

This paper cites Delving into inter-image invariance for unsupervised visual representations.IJCV.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Delving into inter-image invariance for unsupervised visual representations.IJCV

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.386791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:dcfe386058d244bef7768153d698d41e2f41882e3c70419516ea2a08e79d3d3f

Observation 9775d78e-cf60-4eda-bf4f-b5f9f2cfe3fb · outbound

This paper cites Masked frequency modeling for self-supervised visual pre-training.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Masked frequency modeling for self-supervised visual pre-training

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.441337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:38a704386c653b0091ca9a1bd0d9b2a0f03aaefb2b5cb146b768894bfff141dd

Observation 98d290cd-b7d1-4857-a35b-5c5e73eca5f4 · outbound

This paper cites Depth any- thing v2.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Depth any- thing v2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.449235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:d3baf3cc3d0b3129498f9948b74276018b92dd9833047be0fd28442578add4de

Observation af98848a-06ea-4dda-8b6c-663bf307169b · outbound

This paper cites How to evaluate the generalization of detection? a bench- mark for comprehensive open-vocabulary detection.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models How to evaluate the generalization of detection? a bench- mark for comprehensive open-vocabulary detection

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.425312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:6fdf2cd1794953fcdaf0d6c61bc119dacec1e7219348182543994f297dc72d38

Observation 37c41524-cf77-4589-b7b7-e8f28b166e66 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.832543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:cf688cafb51f5d0d61269b7df9442c33f70a596639090622d90a0024ce96f475

Observation 013d3648-2590-407f-9943-9e8bc8b7be78 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.633735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:53f6d2f5e5c40ac38809d04568cce49fb8f450c3a8bcbdd7050e2114caa86fd1

Observation 82fd2cc8-b1da-4349-97f8-e1b19fe59164 · outbound

This paper cites VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.628351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:cc1ed1cac424497564beb7b39398ad135a7b7bc74eebeab48c145e3204582df8

Observation 7897cae0-cbe9-414b-a511-49f6176688a4 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.429163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:10f2b010ae6cfa6afcc048d4a7cc6a3102ed801db3e77baf49aefd728931b336

Observation 785a44ac-6f1c-4d8a-8882-d8bc0f501496 · outbound

This paper cites Sigmoid loss for language image pre-training.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Sigmoid loss for language image pre-training

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.433096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:91cdb015eaad446367b2c861407e52facd71728a55264ecf05383182994dd04a

Observation 66c10490-5f20-48f6-a458-0ac81f389d47 · outbound

This paper cites Online deep clustering for unsupervised representation learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Online deep clustering for unsupervised representation learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.437208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:4887f683d34ed4e4bccd14f6b776405ca05d380d6263201a9e4bf829cffbe14e

Observation 18f68011-3757-4d0c-9051-27a9782c586f · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.413172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:e227640c83ae0bb315aee14d520a03b392f3c09eb5a397f27a6833cb41af2415

Observation 174fe4a7-f5dd-4cbe-b0d1-fcb98e2163d9 · outbound

This paper cites Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenar- ios that are difficult for humans? InICLR.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenar- ios that are difficult for humans? InICLR

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.417301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:d0bf89b81ff7e8dd4d75abfdc244b12bb6c44cf7f27420f27cc35043006517dc

Observation 43e8db37-78f0-47b6-8ad9-cfb67458cd7c · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:42:57.181531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:17c8c325d596a3486f39db75c5667910f24a969ef4c8684696509b01b383681e

Observation c7adc2e5-d8fd-4fb6-a0af-f7b9fd059301 · outbound

This paper cites ibot: Image bert pre-training with online tokenizer.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models ibot: Image bert pre-training with online tokenizer

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T09:37:50.421382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:316f5026cfaa46e0bd3dc58bea6005ad1e904f375a1b025c47b3f95bed0e9bbe

Observation 8811d8eb-aa12-4f9c-971b-d2fd7cab62af · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:36:08.610800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:c7f0e9df1d82f33779699c36ec652e1c8517fe30a83f8cb67ba2bbfdb8b09329

Observation c4efbcf1-98a0-4bfa-9c6e-6c4dd460d0ba · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 88

Resolution
malformed identifier
local_arxiv, observed 2026-05-11T13:36:08.587864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:e5d5cf212bc893ab8d4d6b1a84199df0be2ac65528a9e229f1547d9049fee099

Pith citing papers

Observation fd97243f-5c04-4656-9560-48fe1ac9104e · inbound

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement cites this paper.

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T13:56:46.759454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T13:56:43.645518Z digest=sha256:090c2b4f88876aa05187b0dc274d381595c6ff9573933fbb9e1194d9a7f7794f