Pith. sign in

Paper Citation Record · LEDGER

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 9 inbound Pith citation observations for arXiv:2510.12784.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.12784 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:54:26.490129Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T07:03:12.753735Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:39:51.508184Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13b77b4b-7b0e-436d-9ec5-675cc5d63b2a · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:23.856369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:23.856369Z digest=sha256:dac497f3beeed7ae097ff6e5ad8c07f73d943ab2f356100b9e8b31156a2ba5d8

Observation 532963aa-c400-4198-8e4f-e39c5f12a212 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.057136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.057136Z digest=sha256:b6e89d4c6b08066cbccd6781577824780ffc4d1dd1bb24fdc94fe3e6ce60bef5

Observation 947d853c-d5f8-4a92-aaad-6d0b9e6cc133 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.131408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.131408Z digest=sha256:787dbdb6758cd8352de04a51ee459bcd909d7e09d5fa07db623ec0f79a6e7d9c

Observation c317703b-34b2-4365-81ba-33f4da95b3ef · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.239016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.239016Z digest=sha256:e410683e614e9c6e357b79b3529fdf1b5b7461b3b63508b5dce7407a17a03b28

Observation 426f6dc4-ea4d-4757-936b-199b4552a7d2 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.247164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.247164Z digest=sha256:d8b8a31eecada09f6a0fe9350bf22dad135f90f2f18656eac51a4d0a39006907

Observation a9b203f2-4f05-4091-8918-ba685e9797a7 · outbound

This paper cites Classifier-Free Diffusion Guidance.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Classifier-Free Diffusion Guidance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.551754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.551754Z digest=sha256:694c38eae9ed0417b6420da95dca90ec32da312d09fb2c4c32c0652305e08b3e

Observation f6930a8a-66d4-4f86-9732-625c7dd4a24e · outbound

This paper cites HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:25.111416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:25.111416Z digest=sha256:f4ecede8fa3ff21613443e2e0ef21d38b8a394bdc63882d05bba1ed33b38f86b

Observation 17a19361-0ea5-4e8b-9ff8-16d408649dd5 · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:25.305850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:25.305850Z digest=sha256:3e02e8a9c36ed6ad073f09228714b46b7286cf06f9a6b2c82fdc1838b2ccdde2

Observation 73fba6df-8a41-49d9-a88c-9f7610cd1f28 · outbound

This paper cites Langbridge: Interpreting image as a combination of language embeddings.arXiv preprint arXiv:2503.19404,.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Langbridge: Interpreting image as a combination of language embeddings.arXiv preprint arXiv:2503.19404,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:25.476201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:25.476201Z digest=sha256:4db0e9c2ca879afffd69190e86caf91bbf66fe3c22ef6d87bae9b7122a9c3750

Observation 6fc03e94-81fb-42c7-a47d-c53ada4e7b1d · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:25.696314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:25.696314Z digest=sha256:ff29a4168af65197b8cad66f1880b3076c775aad3bd1e545d8ba679fe4523248

Observation b0332082-7bb8-4834-b58e-31bdf2308bff · outbound

This paper cites Decoupled Weight Decay Regularization.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Decoupled Weight Decay Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:25.973674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:25.973674Z digest=sha256:3bc491deb73badc8f1bab20b10e0bb59b1fcfcdb888e3cbe64e2b096e6d62fb0

Observation 1b38343b-63f0-4da6-9d9e-c2958e2738fb · outbound

This paper cites UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.172394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.172394Z digest=sha256:bf1cd3480d1eae089b9e75bb233a2fc8996707a16b31f9df8026fd02401afef2

Observation 69d936c1-0a4b-4c8d-a1ed-c26e1e006dbd · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models A Theory on Adam Instability in Large-Scale Machine Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.229865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.229865Z digest=sha256:e1dbb9bda411307b306ebb8a3d67f7fcce8ce60e2e24dcecd4c87981adad1361

Observation c46305d4-fff1-43f3-9648-b6231f0a8772 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.236524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.236524Z digest=sha256:9f7a1de9b60f9b8cf1d308e7554a934e687ed0e71f7f123a3e07ee7a884e37a6

Observation 43ca6ec3-72ee-4f1d-be10-ee0266735696 · outbound

This paper cites Transfer between Modalities with MetaQueries.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Transfer between Modalities with MetaQueries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.356247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.356247Z digest=sha256:57df548a44dd4833fbaee6211dbf2bb145418dfd3eac5711af376986578ba26b

Observation 527a0908-9afd-40e7-91e4-9cdb4c2ca112 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.449060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.449060Z digest=sha256:02fe8ee15e99a148da7d3331d59c9a3ba478c5a3487ccca8ee753968af4b21c8

Observation 3d150c1d-1de3-42bd-b40a-4a5b6f331dc0 · outbound

This paper cites T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.459094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.459094Z digest=sha256:3a18a0592c2a2a9042abdae14f793783797531bc3adec28a043a16335a81b15f

Observation a631c8ce-52d5-436a-aa24-2d3cd8ab1869 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.461078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.461078Z digest=sha256:2021feb4c6b4987f4870029ca70470286e7af5e4cacda471992ab468d6862642

Observation 4bde2f6d-1143-4f64-9030-7ec9e531620b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.463571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.463571Z digest=sha256:8c2bfa748f4eda145c1d9ead4f320318ea7b610a8aea8beb9eb41d1ebe3342fe

Observation 9279a138-2b4f-4018-bcff-553a3dc28fcc · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.465656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.465656Z digest=sha256:500f41f47c6a8238f5607b4b3febe6751ff2a19c89b9e5a7870c30505d765997

Observation 9c34c403-4ceb-4df0-a2ec-b05a656ffa54 · outbound

This paper cites Reconstructive Visual Instruction Tuning.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.467952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.467952Z digest=sha256:ab583311735797303d8f495a548a7c5ed654e408c2379dae8ee20a2ba39311c7

Observation 4bac751c-412b-4afe-b8f8-d54e3c22c915 · outbound

This paper cites Qwen-Image Technical Report.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Qwen-Image Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.470131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.470131Z digest=sha256:fcf112385aa4c342ab9615f9099ff5c40afd8c7e08c2e1325f76072fd28803a1

Observation a54366fb-566a-446d-90c4-5088e15327a5 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.472229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.472229Z digest=sha256:62a53dae712cdf6d3ebb19e8635d5d1f37d492e4f9c0fa184a8cf7aa68b5219f

Observation fd171dc8-e761-4fe5-81e1-ac08122995db · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Show-o2: Improved Native Unified Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.474337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.474337Z digest=sha256:6344f8a282bc4bd6b01a017ec254ca1cc319aae5df8120ba52ecc5f54a6c35e6

Observation 8afffa9f-1194-44f1-b8fb-40e94f36eee3 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models DanceGRPO: Unleashing GRPO on Visual Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.476843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.476843Z digest=sha256:3f417a2e49c405674084ad75e97dbe428f774ad8d0b915466b0f446d063b71dc

Observation 6373a76c-d776-48cb-b9c9-e52ee0f9deef · outbound

This paper cites AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.479280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.479280Z digest=sha256:a0a6fb3526e0e06b79401c398045b04c25d76fd365f5430086e12f004ccc2a52

Observation 3b9d40d5-b6bd-490f-9d3f-24a5d8907f3b · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.483487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.483487Z digest=sha256:91e34152910f0ce61d7a47339b1518fd749385e60e34fce5a08205951e75b16b

Observation 97895aa8-09d5-4beb-8f1d-93c72b8a4cfc · outbound

This paper cites Drawing inspiration from (Molybog et al., 2023), we set the epsilon value to1.0×10 −15 to mitigate loss spikes.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Drawing inspiration from (Molybog et al., 2023), we set the epsilon value to1.0×10 −15 to mitigate loss spikes

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.485681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.485681Z digest=sha256:6bb4db1f88c271ae5744470503c1bbb0f3347da2a1e2e7235e07d78405467391

Observation 36ba23e2-6abe-4e01-8f55-e3824ca83694 · outbound

This paper cites an unresolved cited work.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.488117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.488117Z digest=sha256:e4cf5cc4dbb3890489fa58ba40c73894872867123cd950858fb79acc93fb0137

Observation 884fcefc-72a3-4dee-bc33-1b25b51dd186 · outbound

This paper cites {original_prompt}.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models {original_prompt}

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.490129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.490129Z digest=sha256:f9ffb6992c74d8f814690c922e9cb9889e7c11edf12fc218ed75634ede713b9c

Observation 514b78db-af38-44dc-9eff-39a5e4047e4d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.405328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.405328Z digest=sha256:36baa922832cda1467e252cb12ea5453a48c1ed0fadeb1fc3d3c15d0b516f32b

Observation 0e88a398-161a-422f-9793-868b228fd1dc · outbound

This paper cites GLU Variants Improve Transformer.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GLU Variants Improve Transformer

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.434173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.434173Z digest=sha256:badbeddfc3cf07fad57bb3e4129f3222fd2f47c3dfa4d92a7419ba3f0993ce64

Observation b5313f9c-29ab-4476-ab86-95b5e6ea8956 · outbound

This paper cites Training Compute-Optimal Large Language Models.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Training Compute-Optimal Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.763296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.763296Z digest=sha256:3bf3d40b3882bd910f7a7af8df59df5fc0ff4968deb363c6ff529e4f26e49d56

Observation caa6bf72-8cd2-47eb-9a59-4c222b40aa02 · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.184568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.184568Z digest=sha256:081c1d3a186452939044ccf54edbdced375d1ac3bd2ec67c31300ecb2f4702fc

Observation 7b72bc66-ff7d-4a96-9026-21567ceda278 · outbound

This paper cites SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.937137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.937137Z digest=sha256:e2887a7facce0b773a4ee18788ea5d34123f5e69c9d73f54226b90664ada0c85

Observation 3c77074f-a62d-4241-ad3b-60d46aa71d6f · outbound

This paper cites HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:23.900864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:23.900864Z digest=sha256:69fa3865e9e397f7239fb5b6e710c69bac8bfe03fe20c30f28cbdf3fbf382ccf

Observation 01701d6b-77cd-40d3-8742-3dc98c7057ec · outbound

This paper cites Qwen2.5-VL Technical Report.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Qwen2.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:23.983819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:23.983819Z digest=sha256:1cb583323156babec7f70a0eeeade55cbdb49e238293c23d27c5163a60f83009

Observation ad0c698d-34f5-4faf-b73b-d71726bf7048 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.241226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.241226Z digest=sha256:20b40afbfab463247f5ba2cfdc111cd3ec4889da415f59eae454eeb6f3ac4fac

Pith citing papers

Observation 4cf4f63b-3c5e-4d8a-ab65-2a4809a517b3 · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:10.489172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:b5fe2432a9e25442e53953e588aa3a5c27a99aa26c77ee6bd0d416af37719d19

Observation 50317f72-eb0a-44f4-a0d9-edf4bd10776d · inbound

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback cites this paper.

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:10.489172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T18:01:19.748677Z digest=sha256:c31e004d8cf8d00bd0b32a1a5b7796e77a03059e745372aad112770a5189adab

Observation 4310b37d-95b0-4a97-a2c1-5e300df3591e · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:12.753735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:12.753735Z digest=sha256:871f00b2df201c9c36b00c374fb36b69d14a43622d9209433496e79be36be196

Observation 951fc723-3b58-4cf2-aa41-eccb5dad1761 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:10.489172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:bf0c6afd7c4106af0fb8ceecb7dc499a5ef3dd9851d52967dda50f53609c3cfa

Observation e1b01783-1022-4eb3-b66e-130f2affe4b0 · inbound

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement cites this paper.

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:10.489172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:08:57.793923Z digest=sha256:6f7eb4f3d0793a66097f3ad1f9ddecacefcb09fb010be504335d0f28de8ca591

Observation 23c978e2-9b35-48b7-a8ae-758dbddb341a · inbound

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation cites this paper.

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:16:48.490594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:07:06.056441Z digest=sha256:7e61076ecad31682a52fda6e299a1ea1a655b36a6d9000c0f75c63b0ed51f8fe

Observation 3c66c9a6-e1d5-434b-840b-80d287bbf7af · inbound

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards cites this paper.

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:39:51.510296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:58:15.891214Z digest=sha256:ea7d5ff57ec1322c6c526efbee8765d1135b6cd4aacbec1b0cf2baf8a8f0878b

Observation 9ae340f2-453f-444f-8f11-073b87cde87f · inbound

Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation cites this paper.

Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:24:41.122344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T05:55:54.345160Z digest=sha256:4dca1cc4964b421b924ea0d782f3283eee473be8233b75c1e98770a3d5b6a0d4

Observation a544d83d-0cbe-4b99-aa04-ec25b715ca85 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:21.368364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:21.368364Z digest=sha256:d3949f5b4d017f2439bbaaec8b5ca6ee7a458869ddf075f57fee6b4fdbad7565