Pith. sign in

Paper Citation Record · LEDGER

Meta-CoT: Enhancing Granularity and Generalization in Image Editing

As of 11 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 3 inbound Pith citation observations for arXiv:2604.24625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24625 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:30:28.636915Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:23:41.271521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:28:58.249219Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact51
  • verified fuzzy32
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2874846a-c909-4187-bb09-5518e054697c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.360977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:4e2f065efcffccb3d6d5722ff983f607c0489d3edc4f7eec0ee5a8b2724e9058

Observation a78a23f0-fd8a-41dd-a087-df2d0e9e740c · outbound

This paper cites Qwen2.5-VL Technical Report.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.383070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:59c4ca0e95ab2cfda693da9685ab9ca81a3629d2dddf473f9c4570ee65baada4

Observation 9451513a-ac71-49f7-83b3-8e578c0c2cba · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing In- structpix2pix: Learning to follow image editing instructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.747717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:5f42522b6655d4b864a86039864599b74f3dea87a0b9b99ab2d4b1bfc9de93c2

Observation c64ff326-3b9b-4328-a666-044ec560aa0d · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.216281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:a67ecbd1426d17078c66c43dbf6b6d7d916701f708fe95aa28e62185b1e99892

Observation abe6e895-6ab0-4119-bfb8-9dfd70fff197 · outbound

This paper cites Blip3o-next: Next frontier of native image generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Blip3o-next: Next frontier of native image generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.222593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:63aa489f9a69a2f01898391cbeb6736015911c68c2ffe10b35d2ce085a048f5f

Observation e8ea0ca8-cd9e-45c2-9baf-38e0b3f0acbf · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:14.055858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:9eefb402f5c04ca9d1b63423a5167046b12ad9697a4d5fc95d3fcbc0671a587c

Observation 38fdb81e-6299-4b00-bdb2-00d92e0fe034 · outbound

This paper cites ChatUMM: Robust Context Tracking for Conversational Interleaved Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-02T04:04:28.060339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3fdaf746bfca6b09dbd5a9a9e5f92aab76c98d4507637eb3d36df85cde466615

Observation a67b3420-a60c-476f-94ec-7d97852d5a32 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Emerging Properties in Unified Multimodal Pretraining

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:15.005174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:d62ea13318fc45abc3d20e4dbb8b7b9ebc6bd50a17ef52b8226a932e1c19a22d

Observation e0f0ab6a-cd87-46e2-989e-2f305261a917 · outbound

This paper cites Dreamllm: Synergistic multimodal com- prehension and creation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Dreamllm: Synergistic multimodal com- prehension and creation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.721789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e9193f1e0ec98cd68e62354df052722e648e6454bc97a633bc5edb17939a03d8

Observation e500cf13-84a6-4c94-95e4-5c45e0173aef · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:13.279370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:9c528a411ccc749e99db5b4dd3ea2aac86bb505ffd7b01a1d46e59b4b1967f39

Observation c251ec09-3171-4a66-8787-bfef63b28c40 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.274947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:39f770814682114cd38a170c1b18aefeec86616d5b24cd9e682b2c8db86af37c

Observation f704a458-45c9-43e1-8da5-c1b1fe9776f4 · outbound

This paper cites ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.382365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:33a6fe9d293c803391f0118a5cc059070e95f93ba73757eef106065ad9bdc1fa

Observation 7cd4543e-76bc-4aef-82f1-3d065e399684 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.313335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:86765f19005ea797475c95968a55a1062070f2dd255a6ff276a8d1e2effad8b0

Observation 2e74d709-3e64-4ee9-9ac0-419bbc4f41fb · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Visual program- ming: Compositional visual reasoning without training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.731617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:be11d7b0525459a867d44bc23366b82e8981b94707624729290c464d4b415dbc

Observation 99275ec5-d445-46e7-863e-b3cad79f3f72 · outbound

This paper cites Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:13.486416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ef5fa19d17017d5eb67f0045cdcef303268585359f6888b3cab6019e355b3ac1

Observation a5a5a9e4-34e3-42e7-9d68-fa2c3498179d · outbound

This paper cites Freeedit: Mask-free reference-based image editing with multi-modal instruction.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Freeedit: Mask-free reference-based image editing with multi-modal instruction.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.738145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:99fe44ce771dae35198e57a487be7dbcedd4bb2710e0045cf4d4b71c892e1892

Observation 88dc8d79-2fae-4b63-a436-cc0ab1b8eb68 · outbound

This paper cites Re-align: Structured reasoning-guided alignment for in-context image generation and editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Re-align: Structured reasoning-guided alignment for in-context image generation and editing

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.157170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8573042c90fbc66fe118a4cbdad1c2acde754a1364fec6b85bb8ea1d9deb1fb7

Observation f312a434-02c4-4b04-845b-00bc95d56784 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.NeurIPS, 37:139348–139379.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.NeurIPS, 37:139348–139379

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.753884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3b3d011f3c30e81d9811d21dc1e857d3d5560ef2c3cd34b215be5ac11197e6bf

Observation e3172cd7-0357-4060-9b32-6c9124654e42 · outbound

This paper cites Image Editing As Programs with Diffusion Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Image Editing As Programs with Diffusion Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.360685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:737172cd823a079f5845a8e6a73a81a4249bc27b326b84e7a34145abad8a1ae3

Observation f4689a43-976b-4aa1-9c08-4105a9fbbe3f · outbound

This paper cites Large Language Models Can Self-Improve.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Large Language Models Can Self-Improve

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:00:48.316458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:742a651102e20bd4003dd7fbd37268a0acc37d45a76b65ab99bc1a01039a51d5

Observation 78c1b218-e674-4b05-b64c-8811eeaa877c · outbound

This paper cites Interleaving Reasoning for Better Text-to-Image Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Interleaving Reasoning for Better Text-to-Image Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ac9ce4b9b6e853bd9290ecdd5994290ee886dada2c27c1e0779bcc1e7d686afa

Observation a73b6f31-13d3-4787-a20f-47c5863a97e3 · outbound

This paper cites Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.750885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8030aeebdc102f60f1fc91140f534c3986e959324ef4caf412ab1006ac25c2d0

Observation 660225a3-82dc-469d-bbf7-5c2231f9b071 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.265561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:1ea42cb20288a97cb8f98c917874614bcd178af650502e81717a8c86a22a5b6a

Observation 81b3d0eb-df2e-4e77-ac49-c87c051354fd · outbound

This paper cites an unresolved cited work.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:53:02.724507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6ab81a2e549fc9d5837c3c891975f514566cdf7ee5af18e42aced0b360e70ecb

Observation d48f321e-3acf-45a1-943f-faf5e2b1b736 · outbound

This paper cites VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:19.200450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:56d7773fb3cc6c9715e61d8377cbf54d59267caeeb7d88b86ab5f81307a39890

Observation d306eb2f-bc74-4aaf-a228-a81bff10a227 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.248810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f4a78104717cc103b287f6b6426ae8781b3d4dffe6fe1b4961f7ff5e3fe039b3

Observation 872c16db-089b-4603-a488-9ee45b070041 · outbound

This paper cites Prometheus-vision: Vision-language model as a judge for fine-grained evaluation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.727850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:3cbf0977e8d1503bfafd5b617eff6ed329cec0003c689018a61483c71893ba55

Observation f929c851-287c-483d-a507-922dd376e1a7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.168309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:81313e1da3da98066a4be80add88ee555d84c807006e3f783267a63da1e6286b

Observation 0724f602-58d2-42b9-9adf-cf1c64ca76c1 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.995208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f2da3ca8a24088661fafd37568f9972e31ac7ae2248658dcdd6e8e741f5ae603

Observation 2bdb0531-36c5-4ea3-8f4b-b7963a30f91e · outbound

This paper cites ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:12.593382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7d1c367c12b47edb971e03f5bbd4b4286979c32a1b4d2c52f1f86a483fec7828

Observation e1df8cdf-e095-434f-ae49-2dac8b1e6f08 · outbound

This paper cites Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.734711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e266c6833ee7b4654e895983cbf71ce8b39e89517a2a801d5f3f3b18a21949bd

Observation 756f4290-0c5c-4196-bd51-fc05f3e49e5e · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:45.158778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:5fc5b4e023bede9bae8627ec0a184e1e3981de1925dbcb2c957522a7202ca460

Observation e7352c77-fad4-4aaf-ad4c-57d751daddd9 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:26c5a7e5ad10233fcae9e798caa009360605cffbd71f30554b2c6bc0146a7ad5

Observation cc8a5292-4147-40e5-866f-cbb95b73c7be · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.304691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:d8f832191a3b2cfe4d462aa7e685c6d9dabcc3c23ad0824b25486c9d7990cba4

Observation 7c33a7aa-d548-43d8-bbef-b586cbd9fb8f · outbound

This paper cites JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.240606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6396825c1adb4085584cde70409c1a5b84b4a2b3fb889118df7ec488d5b9c19d

Observation 55bda6c0-36bc-4bf9-9f8e-95778907e6dd · outbound

This paper cites JarvisEvo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing JarvisEvo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.178955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7b34700e1034615c3d12a05add6da88020b19619c24964ac4cd8305f7d0f85e2

Observation 10895829-856d-46f2-9f14-1a283cc64d6f · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Flow-GRPO: Training Flow Matching Models via Online RL

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:15.121348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ebf6c3cff6c4a4138a32b1e5a28ee3a11a8623e00accc1ba6b8e2a47fb55ba1b

Observation 21440f07-a99d-4170-b9e5-072bbb325427 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Step1X-Edit: A Practical Framework for General Image Editing

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.319109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:caa94575ef37914b22c55e6ba2aae64c101c978e07325461314b667c6d9e2a97

Observation c0f9f15c-61ab-4142-8d71-357f084edb35 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.741290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7b4cd92f98f1f641abc39e4d42a968ea43d14df25d89bb18fca0b74e4bcf1bdb

Observation f6bb620b-1011-4536-b087-b4b9e68fec27 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35: 2507–2521.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35: 2507–2521

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.756900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:17b2e24096de9e2ffca6e6e39269efd3c3f2964c6089276342fb4cb996cc7a99

Observation 5cdc5a57-7228-4b44-8af1-82bd98dad454 · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.759771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:4c92b02f67499c974a543bee63178e4c6d67ac6e7b137b8863386851febd4792

Observation 60008d92-5c65-45e4-8f25-fa10b455564d · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:12.347080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:b707c2e427d1f1d71c91b377b7d1fc4e834f54e3e0204351aa0332e6446c997e

Observation 67ca32ca-fb9e-4c6e-9a85-b183b026f6d1 · outbound

This paper cites Fastvmt: Eliminat- ing redundancy in video motion transfer.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Fastvmt: Eliminat- ing redundancy in video motion transfer

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:16.183811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ecda1ba8ae41ce7c3f2ec4a14f8e2437f07156d6c5c61027e916f18cb2100e84

Observation b716aeb6-631a-453f-a487-15391d81d25b · outbound

This paper cites Gpt-image-1.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Gpt-image-1

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.744311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:924c5b2057b94ff08f36554795d0b632b8329510f497605d19bc10da2655c5bd

Observation 0594f594-434e-41ea-afa6-31cef20c1094 · outbound

This paper cites Introducing 4o image generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Introducing 4o image generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.702266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f323242c966784a9d019cfe55e4f68c53da1929710394182b34694ca3f9890dc

Observation 1c5628e3-af67-4ad0-ad82-63288334c44e · outbound

This paper cites Transfer between Modalities with MetaQueries.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Transfer between Modalities with MetaQueries

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:49:23.311693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ac1e2b09207ccd1339c6af465b46e90b4cd1817d042230a65774222a41f09c18

Observation be94396a-a4d1-46cd-a52a-af3776768c43 · outbound

This paper cites Scalable diffusion models with transformers.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Scalable diffusion models with transformers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.685938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:401dfb5ce1525adae3cc4f8e61b7dc1e3e28fecccc56ac359b3694e700d41c8e

Observation fec4d6af-7265-4a0e-a833-957939712b01 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.367534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:20ac614376dcfc9f8a1b508985eef7d7b020f897d7929556da70268109c1b5ca

Observation 9b98ae0b-e1a8-43d9-9f6d-162ce98e9c99 · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.NeurIPS, 35:25278– 25294.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Laion-5b: An open large-scale dataset for train- ing next generation image-text models.NeurIPS, 35:25278– 25294

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.698904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:211d4b86e80ebab8f5dff3219c1c40f930b29d13bb621ce68ffa812a3b495626

Observation e52a145d-34a0-40e6-8743-b0a0781312ac · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.692609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:50dbf7577d45ab4eb9fb1d0b4ea58f0867797b0906f237e724e6f4221a474977

Observation 3d594a3f-26bd-4e08-956d-66a4a06485f8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.232172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:36333c75a225f64340dc8d2b6ff35d45f06c9196fdf7941af5fa327da03526b0

Observation 66ee8cee-fdc4-41e5-ab3a-66281e56cb1d · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:14.691378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:4cba8ab9ab712b09e10d0473b2bf61d6afebeb8d6c3057f0554a2e8e72fc7184

Observation 161c3c34-404e-4c82-9490-1f4baa4ec993 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Emu: Generative pretraining in multimodality

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.695626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:59e054174659b066a1aac6615d6bf19923fe7d68026b563c5798b870229db910

Observation cc5f570f-630f-453d-84a3-f96aa4ebf832 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.340833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:2236f68c3ecb58b27ccd563041280d6f12555ee76add4bb3576ad36c127835de

Observation 465c08dc-a505-4d6f-9d24-1dc2ffdb1557 · outbound

This paper cites On the estimation of relationships involving qualitative variables.American Journal of Sociology, 76(1): 103–154.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing On the estimation of relationships involving qualitative variables.American Journal of Sociology, 76(1): 103–154

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.661332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:b68d8a6ca3dc14e25e0088db8335a0f234913f1eb4c78d0ef55727ca93f2da39

Observation eceae1de-462f-4abb-8b01-5d30f5df3546 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.882527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:0dfe59d4951341f28efef56892800dac7ada4b3b4f0593801fa74c843c95e379

Observation 9649be60-f329-4e82-94cc-5e60515e225e · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:12.896794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:27d81b57d64ec8de97b2b111339432ee274e5b089506953d5e99f5f91e522b2d

Observation e136563d-1708-4680-9821-b224abf183f7 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Emu3: Next-Token Prediction is All You Need

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.293621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:ef140af2e42b65504fed3b4b672120e9ecfa998b242df45b67ae02697db09c53

Observation 5dc32682-7d7e-4624-a17d-5b97596ce755 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.NeurIPS, 35:24824–24837.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Chain-of-thought prompting elicits reasoning in large lan- guage models.NeurIPS, 35:24824–24837

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.675239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:77f4a9ce5daee491fc96313afdfea995b0ade0f6d35a416cbc7f0ef21aa341d1

Observation 510a65ff-dae2-48c0-a321-0d6264a02ce1 · outbound

This paper cites Janus: Decoupling visual encod- ing for unified multimodal understanding and generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Janus: Decoupling visual encod- ing for unified multimodal understanding and generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.709369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:256484bb5ef97896af407bbac1f976454a1262da4377a73498dee5661dd33fdd

Observation 3b0951c5-58ba-4e25-a937-a4fe2693acdf · outbound

This paper cites Qwen-Image Technical Report.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen-Image Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:15.895704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:5def0eee791c704cfb40004e88fd34bef4a114c3e7c8eb2ac3fe3bd3ca780581

Observation 36dc96e8-3034-440d-9ee5-ea5454ba3ee7 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:12.247983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8c18c793fb8a0710f94dd67f81c9da2c7062dc8ea4a1704a78b1259b92bf7e4a

Observation 8802737e-8f9e-45ec-9e3b-31b192b74c79 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing V?: Guided visual search as a core mechanism in multimodal llms

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.672018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c3b72dbccc89e8edc14091d8d5887cce710d7ca15835c35a09689db3a1f9fac9

Observation 1f2d14b7-c754-4f67-b28e-a1cc6d0abef9 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Next-gpt: Any-to-any multimodal llm

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.682123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:a6adc34a22948c788b76889f53e257b38967b537ffd3b9a7257c58459e22d2d4

Observation 31a92358-c2fc-4b2f-9642-a82f0129ff5f · outbound

This paper cites Omnigen: Unified image genera- tion.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Omnigen: Unified image genera- tion

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.678741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:710d2ccdaa0799731a5334e9dd8208278d6fde9263f75002a3789b64238721c0

Observation fd93135a-5da6-4330-99b7-9a43a7b57548 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:19.375413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6416e286a85a04d17cfafbfa52a93e5fd4730bc28e836b4065bf10ceca95c480

Observation 0ee58263-b97d-40d9-8dfc-66c1e6a3e39a · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Show-o2: Improved Native Unified Multimodal Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.424084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:88123d030215ce85116eb5a15721af65d4023f2cbbe641f6d9ba087f5b62e5c6

Observation 2a62e55b-571b-4b3a-9b9c-189ddd77d5b1 · outbound

This paper cites HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.256048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:2a8ffc9536b96f32dd91608098931b8a905904a12a116b745f4af2122a4885e9

Observation cb10bf85-c4dd-46fb-9be9-edc8b41d8094 · outbound

This paper cites Tag-moe: Task-aware gating for unified generative mixture-of-experts.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Tag-moe: Task-aware gating for unified generative mixture-of-experts

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:14.436382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e312da674a060de274bb18d364b4168810730f54f3a79e202b14ad8d8db28762

Observation 46081f00-a955-4a7e-94dd-dff49f1de582 · outbound

This paper cites Qwen2.5 Technical Report.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Qwen2.5 Technical Report

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:13.357969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:95c15d7774739a88a00241074f69d61aeeaa83b2e19fe3f1ee9993764dcde2e5

Observation 71e30a2b-0cb7-4087-b6ea-f1a9a9f6e029 · outbound

This paper cites Uni-paint: A unified framework for multimodal image inpainting with pretrained diffusion model.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Uni-paint: A unified framework for multimodal image inpainting with pretrained diffusion model

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.705890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:10abdd2f5a69e6f9ce5b3754b7f0761a0012dc8b829711d17cbed19dc658dd54

Observation cc25ebff-3400-45ed-b60b-55baa2014840 · outbound

This paper cites Direct-a-video: Customized video generation with user- directed camera movement and object motion.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Direct-a-video: Customized video generation with user- directed camera movement and object motion

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.664587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f20e6927845cb450ed141a577bff8afb3d68fb6c3759058d7007300f8a076e4f

Observation efde17d6-ad2d-44dd-aaea-9b9f2aa228af · outbound

This paper cites $\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing $\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.354347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:087d8e843076e2f9d003e2a824dd0b0231e0d67e1e8f3055090e212010429995

Observation 1a19edff-0b81-4eb3-a4aa-fea9d1253aa0 · outbound

This paper cites Multimodal rewardbench: Holistic evalu- ation of reward models for vision language models.URL https://api.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Multimodal rewardbench: Holistic evalu- ation of reward models for vision language models.URL https://api

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.750831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:b8b139b0b9e768ff3f975145efc028a62daba49e4ef9ce1d49b5b57877f06579

Observation c0dd0fc4-9bfd-48f1-8795-d705e41db21c · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.693018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:297320dff4a76ebd016da2a1ed50fce3187cde96db981e93c2c8e970e1874a99

Observation 9db8e75f-a664-4a45-8a2f-1bb1886b8f9c · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Anyedit: Mastering unified high-quality image editing for any idea

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.715699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:baa8e1f8fd05eb2ff493487c6aafdb760f93576f1d13bbfde68bbb57b220299d

Observation 3fad433e-784f-4789-9415-48455184384a · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.NeurIPS, 36:31428–31449.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Magicbrush: A manually annotated dataset for instruction- guided image editing.NeurIPS, 36:31428–31449

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.718964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e6ae1667437fe6e28f09270d961494d731a872716bdcf4e4b35b1095367fac5a

Observation 9b5f7015-e7c9-4fe7-8366-46901c619d5a · outbound

This paper cites Logo: A long-form video dataset for group action quality assessment.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Logo: A long-form video dataset for group action quality assessment

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.712401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:7e702063d3a8e4cafcd9fa1bb3651e78b2f6a48359f52e95fdd4100e15cc6c73

Observation d05b10c7-fc7a-4060-9475-ba70dfdfa0ad · outbound

This paper cites Narrative action evaluation with prompt-guided multimodal interaction.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Narrative action evaluation with prompt-guided multimodal interaction

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.689075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:91c2bbf68034cd5d9700c4b21e19f88c5f914a6303068a850a6ad04d3a64de3a

Observation 83242267-11e7-4be0-9e72-e108910d5a22 · outbound

This paper cites Flexiact: Towards flexible action control in heterogeneous scenarios.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Flexiact: Towards flexible action control in heterogeneous scenarios

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.668687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e2925c9138ab92ae7259f85e31f7adb6290ee54e2074813e0169effac66c0a6e

Observation b6533477-c6b4-43e5-8708-eb4dc00684ca · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Multimodal Chain-of-Thought Reasoning in Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c8f591a5341593a3efc284984b79e53e3b5b60a4e0727608bbf217e6469837e3

Observation 413de24d-53a8-4232-a463-30151c361b93 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:07:53.247084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:298bb28fb771a0c20382ff52576acb6b8b9b8524c514d5835ce37e2c7f3b805d

Observation e9a17fd4-de51-4ae6-bd7f-c38be8a334ab · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.NeurIPS, 37:3058–3093.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Ultraedit: Instruction-based fine-grained image editing at scale.NeurIPS, 37:3058–3093

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.654541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6dd0ab1ef93a3e62e2d39a191aee9eccc6f26d6cd1fe5a99582fdc9e5af0c561

Observation 037c69fb-a8b6-40f6-bb6f-ccdf74121747 · outbound

This paper cites Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:13.957691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:9f45aa250f22bf1cb97ade7aadc7501d30933f79a1c5895386395a5ab3a59163

Observation ee8cee85-5f76-40f0-abf5-7c297284993e · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:36199c5784c25d1c6875a3492132f69effac5f165c3adb712f67767f122f79ad

Observation 8efda024-697b-4795-8dd5-fec457353b44 · outbound

This paper cites Kv-edit: Training-free image editing for precise background preservation.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Kv-edit: Training-free image editing for precise background preservation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T20:53:02.657787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:64e86723118a75484eeb3e6bf91b7db6bafce6627acb2110288e3ff5f8e9b4ba

Observation f097a609-7b8b-458a-922a-e6b8fb7312fa · outbound

This paper cites ColorFlow: Retrieval-Augmented Image Sequence Colorization.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing ColorFlow: Retrieval-Augmented Image Sequence Colorization

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.148799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:b4124e5bc1319a4bf84601ffd16ef79337440fba8e6a9a45f99fcd75b87ce736

Pith citing papers

Observation ed7c5895-4410-421c-a7c9-af892091f48f · inbound

Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing cites this paper.

Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:41.150231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T13:37:09.892339Z digest=sha256:97a4b8ddcbd4c1e110b14df81e9b5a9d39e68217e3f899cba250c5c9f672939f

Observation 81170c89-5dcd-4806-b952-1138c72f9bc5 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:56:59.012818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-02T13:56:43.671622Z digest=sha256:6f38ba8615af301bc4b23aa927a661fac898ef85d065db3aa2c1b7b02212d2c3

Observation 2cac8ef2-7c57-4dd2-a656-8f1d4d6fb016 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:28:58.250641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T21:23:41.271521Z digest=sha256:2be9840ed93c7e3bd897dac800fa9167f25ed04a873f910d4b51a3b64ca1277e