Pith. sign in

Paper Citation Record · LEDGER

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.05501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05501 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:23:12.587345Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:12:44.268199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:46:26.764949Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96a77d7f-fa98-41a8-9e07-edf007569884 · outbound

This paper cites Qwen Technical Report.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.478364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.478364Z digest=sha256:e5cf1399958de9803590ac122c80d7dba4108d213ec37e65ae5cc511750124c0

Observation 200d7dc9-908e-44b5-9d48-9045e77bda97 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.481259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.481259Z digest=sha256:b11a69e69f99821e0c5359bef33d0588a099bd8846b8b4eb1b7eab873316cdcd

Observation aa50303b-20ca-4208-8b1d-4821dec5f712 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.483699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.483699Z digest=sha256:58cdc03ecb5d3502c99a0060f18d9004e575d05a16e34e77e14b4a32ec14619e

Observation aa006ef1-21b6-45e0-857e-fbe5e29c86db · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.956200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.485922Z digest=sha256:1d26472c4a238bb98f9585a0e00e1ce7a4d0e35c48b32ff69ed9d92b116d58af

Observation 5b7d0096-7748-4a23-9d69-4ac1d3a6712c · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.488547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.488547Z digest=sha256:693923394bd49f9fa4d690ba69a25ef32fb9c1af32e09d8f8e0ec93edde1d62f

Observation 7ef8ac61-7805-45dd-938f-6a7e989dbefb · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.491082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.491082Z digest=sha256:4edcf4db6c1c77a47c20f309cfd8e6933a772eb6910b24f71c4ba04cd7c0fb03

Observation 8078d485-d69e-4472-96ea-3820844c2e71 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.494017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.494017Z digest=sha256:3eb6e513d6816f58c14c0f5e1a2981ba0403572e2f9c26519990b449003e1efe

Observation cb7208e1-e5af-4525-994d-e3bacfddec81 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.496519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.496519Z digest=sha256:2d7ff860327a3da4ded9582322a24a059aab3eb8c727874bfb107e100b1c317a

Observation d38f3067-beb7-4847-bf6f-4fa8939bcc51 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.498502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.498502Z digest=sha256:1d5f5666f337d3bc2347815a568912bcdd9362455f17a4034fc94c08c75559e5

Observation f2ff5d2d-5d0c-45a8-b643-68d21febdb40 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.500546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.500546Z digest=sha256:155dc44ae7a0d3307726daa4fe1f674df6fc2aeedc4744c103afd7ef74f0863a

Observation d3487a38-dc62-4b86-bca4-cf93de2dbdb1 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.502691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.502691Z digest=sha256:4429be6db7a985281e010924eedc2e3d92c01a86b75ec17c1613cfd949d5f239

Observation 43cabedf-93fb-4bc3-95f8-3054e75a5a97 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.504995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.504995Z digest=sha256:6d90230e8455ffa4e9e5f7834602de1058f3a82dfd34ef5e26e6fb9ae3bc892e

Observation 729a06c6-ce8a-4ff9-b2b8-40ac076a70a3 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.939100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.507148Z digest=sha256:d1af8ba2eb214ba3fd52d336fcb1854e325ca56c1281b54cdaea80dc38e7861b

Observation 335c3dbc-5520-48c4-9bd2-b300b359f201 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.509070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.509070Z digest=sha256:317da6c8e77d547911a48c166daaf51006e693791ebb1689a2d2023df86cfa16

Observation 96b75dea-f716-4ad6-8b6c-2b525eb43572 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.511095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.511095Z digest=sha256:5c857220cb4a84af28daebc0fea1528e16ba207b9b7acb36314544f98e00f665

Observation e2855e21-d9a8-4b73-972d-110d7f90ec7f · outbound

This paper cites Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.513374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.513374Z digest=sha256:823f7d5d9897d4185fb5d7895fdb8e20181a183fa8be95d0c56a6584822093a2

Observation b80772bd-6c4d-4a8a-a5b8-b41a05f3c324 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.515817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.515817Z digest=sha256:524f6695e7f785723d0cebd9aed473465d94b20c7325d0a0a8a572a5526bf1ac

Observation d3300954-a378-45ed-bfd4-cba46133dae8 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.518502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.518502Z digest=sha256:1d60691c57351db208f821e8c18f659b76a4634ece00667f1361a6b20bd327b1

Observation 38bfa069-c6de-4947-9281-2cf7158039a3 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.520489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.520489Z digest=sha256:ed464b012e1f5c8be80a515cc6ff0221872b5ffd6032f4beb9bb600899b396ea

Observation ac842250-4125-431c-9591-e5c9565c852e · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.522483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.522483Z digest=sha256:d1ad2b89d38538104b15bda1b7cbe4c6fd532ee6b5e47e47e944c0a01e83fbf5

Observation 2702fb33-86bf-49bd-aac5-964a85abd772 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.524294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.524294Z digest=sha256:5ed3b504d8ccc8f14ed07685641318be1304830b81c8f70feb24ea57a565923b

Observation a5d08ef2-8fb8-46da-aa02-8e3f1342f3db · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.526429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.526429Z digest=sha256:60b6389de6beb90a1c2722d4ce591249a7707b48b09e8c72241d0d8f71c79270

Observation 103a01cb-0bb3-487f-acbc-489b357ee2d8 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.528420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.528420Z digest=sha256:5ef8339d874db25031cfff4dd14b61cd674810b50e9753f3b6e28164c1862259

Observation b0172b25-e168-43c2-a262-98253bd749ec · outbound

This paper cites Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.530267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.530267Z digest=sha256:d97891658f74bac5c6c4ea455c82716ec08a4dabe66c97f7e661bf38dbd64192

Observation 3b20debb-3f5e-4a67-b117-b03d9d4b5172 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.532475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.532475Z digest=sha256:abe6afc26c007eb66cb988339088f50af8f619b9d2ab85243b8a646c70581118

Observation 9b68318a-909b-42d0-915e-5c207f2ee8c7 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.534533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.534533Z digest=sha256:b900cff18675c06f3e506ecd87349a20d6a0a9cf9f818def1d9293538348cf0d

Observation 9c5eef34-f8f8-4ee5-a798-8c6b7248206b · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.536598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.536598Z digest=sha256:a0bee32a872745db8b8171f1e0849461d8f64dabecf9786aac0099da139095f9

Observation 9ffdc315-6711-481d-8abc-328255e8f321 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.913662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.538467Z digest=sha256:c007e5e785b3eb28276931bb0a297fd4bf45768859b6dc9849d9d45f0654cb37

Observation 1d87d19a-e1ae-4303-9f62-5672c95aadb0 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.907359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.540420Z digest=sha256:e9d5048ab45f019c02cf65fa08a0cb1f209eea6b7bbec5599b1bb7c798d5b4e5

Observation 7332d79e-57f8-4174-85d8-8b70c25bb5a2 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.901377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.542459Z digest=sha256:7b4a7fc8dfa673fb8cdf42b46d8b71d70478b1bae411bd823ab305390fd74af1

Observation 9b2e2319-32ed-461f-b9e9-6439c909aa14 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.895261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.544337Z digest=sha256:a6f1a4bcabc2586b5009e82c07bec846949c080f3e39fccbdb4a3867fa855ec9

Observation 2774d770-059e-47a3-b1ad-0b65c07fcf89 · outbound

This paper cites Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.546258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.546258Z digest=sha256:fb021af11fe75659934febd596801640172bde0bb16b485055cc2c0c11197a5c

Observation ba6c5649-32ce-4b33-8167-eb10939e377a · outbound

This paper cites Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.548885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.548885Z digest=sha256:a51b70a2762d3ae78ca86f88c587f5fd8403c5a96fb9d8a26643ed63242f9d57

Observation dca5abc9-c343-408a-8b5e-7b94d32c7641 · outbound

This paper cites Auto-Encoding Morph-Tokens for Multimodal LLM.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Auto-Encoding Morph-Tokens for Multimodal LLM

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.551112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.551112Z digest=sha256:236f383cfeb06313b807ad0f1b7f520a32ac4d1efcb64fabad197965bc566e2e

Observation 847cb11b-3ecf-415b-86f8-11836b42ae8c · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.553747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.553747Z digest=sha256:44a490b6484a8f0535f5ccc810ca5599edc1917aa618ab281c5bb7a2d269b85d

Observation c139af4a-b154-427e-9360-3ae217bb71dd · outbound

This paper cites STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.555613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.555613Z digest=sha256:f9498e5185726f2d45661f44d8f1508c3403504325c4d8c537ec335e7613c262

Observation d5ab22ed-e6c7-4935-acc6-2f424e9e5885 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.558043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.558043Z digest=sha256:f3ed7076049cd442416338b10a2eb8afb6b8df27a9ca43108034d807de321eb7

Observation 80576249-4419-4000-b45b-25bfa0c04620 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.560144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.560144Z digest=sha256:10d290d56f374ca449ba2c34e5b487e5672604aab683747b112679836e4e4550

Observation d0eaa143-c41d-4162-9181-63929e916336 · outbound

This paper cites Exploring Bias in over 100 Text-to-Image Generative Models.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Exploring Bias in over 100 Text-to-Image Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.562389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.562389Z digest=sha256:61489f96d9f24fffcee991f18bc79edcc317bdb70412d6c334c5d3d54fcf4265

Observation 4577579a-4ee3-4c6b-a3c7-d3abc8a12f2a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Emu3: Next-Token Prediction is All You Need

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.564317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.564317Z digest=sha256:cb2ea415db166fe382af3d80b1be09eb52bc546247b756cf8fcba919b3f365a6

Observation 6f51ef8b-170d-4fd7-9587-a084cd6aca56 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.566657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.566657Z digest=sha256:72f1ab43bcc5d91f674919948e6301e6ccc0f0330216fe489f54616d30e09b2c

Observation 06cef3f4-e53e-4366-a48a-ab0590a4e052 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.568772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.568772Z digest=sha256:75df909d5d4fa5e249b9e2d6da582660a8d0b1e0811e3cc006bd5234fb5d4393

Observation cc0bb740-0c9c-45ab-97af-f5ac02e04aae · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.571060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.571060Z digest=sha256:d256dc73a6544edf38fa9ad6474d1152cdcc510e8ff623924daee707acac0c61

Observation 002e6126-6d0f-4cb9-9fd6-a1e4c5f649ed · outbound

This paper cites SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.573432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.573432Z digest=sha256:5c7ffe7f98ea9b4708dbd296877371815a30cc07722efeb0194657f280b73e5a

Observation f5bfb895-cb14-434c-94c6-f3ca527b0271 · outbound

This paper cites AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.575466Z digest=sha256:854917470b6234dba71cbf9a93c9886234d8f0f8729a215b7a3526d7584b3008

Observation b21326d8-46de-41c1-a09d-972fc86c1262 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.888567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.578159Z digest=sha256:08254a47f4fa1853a13442fda3c930874923c7036e478cdf82803bb87b81ec84

Observation 806798e9-e5f9-4bf7-bf90-b96e1a959eaa · outbound

This paper cites Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:23:12.617976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.580211Z digest=sha256:982de62fa21ab3c670d4187e80b7ed842207cc87b8375637228b4c36c076f81d

Observation f2de3076-eb6e-40b3-a718-a37c94f2152c · outbound

This paper cites VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.582550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.582550Z digest=sha256:817523c498b288e01bd843fc20096918e9bddded623fea70d7dc078e59c53586

Observation f1cca2d8-79c8-4148-a7e6-e175d8903d09 · outbound

This paper cites online" 'onlinestring :=.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL online" 'onlinestring :=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.584717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.584717Z digest=sha256:1e328eccbe6aecee52e38199e605bf6f72740c311cb86d3f5621af2d48514042

Observation 92668bde-dfec-4198-a361-bb225f2e4309 · outbound

This paper cites write newline.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL write newline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.587345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.587345Z digest=sha256:d0139e2ccae4c07f42248110c186a59e22eacab62fef7e9b471a2346e4532253

Pith citing papers

Observation c012480a-14d9-4fdf-8d47-e429a5d16559 · inbound

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs cites this paper.

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T05:12:44.268199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:12:44.268199Z digest=sha256:70ba28faa7194fc4b6793b636e8bca901e0efe4ceec3ebb74831ca7a11002c5a

Observation 334b2604-7141-46d0-96bc-011d1b7f2b03 · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.767451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:03fc87499cbcc254057cc32ea3711252d4803aefdf592e67573d6d4b587d7dc0