Pith. sign in

Paper Citation Record · LEDGER

LongCat-Image Technical Report

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 59 inbound Pith citation observations for arXiv:2512.07584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.07584 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T08:04:12.949075Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:38.987981Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:09:57.573753Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact28
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23dfbea6-8b83-493a-8bf9-53c6c78b86ca · outbound

This paper cites Qwen-Image Technical Report.

LongCat-Image Technical Report Qwen-Image Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.966785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:8aff6d68983b2886629be369e84bb866f2fd8b7feb915e52593cb970aa80b89d

Observation 5d91a54c-1b76-4027-90f3-7dc99226459a · outbound

This paper cites Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model.

LongCat-Image Technical Report Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T08:27:36.473449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:f2b9c5a90c0b5578b99a56abae2416d5d81b417a51d61ffd52fa63dea906d918

Observation 93e84043-04b9-4e44-9b3e-8a8b8f85733a · outbound

This paper cites Seedream 3.0 Technical Report.

LongCat-Image Technical Report Seedream 3.0 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.974044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:f76da2e14c1ef237fcbc96f7c2631ff52aaf025d7a39f0715b774b1135d8cd88

Observation 17934025-b13e-4c04-8333-dd3f8fae2ef6 · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

LongCat-Image Technical Report Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.977544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:d4d8d6cb3e88b69d6568ac1521f1fa161f9a51d91b64a4c9b24ec7bf6f963b80

Observation 14ba40c0-04fa-4daa-81c9-fbd6ce3c0076 · outbound

This paper cites Image Editing with Diffusion Models: A Survey.

LongCat-Image Technical Report Image Editing with Diffusion Models: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:12.981186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:a1ae4b84ad61629f16f308c33b7d20b055a53ab4018cd350d6f3e63db58859d3

Observation ae853c10-6b92-432f-81b0-d340145314f6 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

LongCat-Image Technical Report HunyuanImage 3.0 Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.984226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:b660de5698c8d1eeef6ff46b70a92174816a935a997e5df44415b8bd997d1e4a

Observation 793e525d-135a-4d65-873b-592e3d4f7f4d · outbound

This paper cites Qwen2.5-VL Technical Report.

LongCat-Image Technical Report Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.987492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:cce452df5acebb428d85ce204318ecc9454696e07c414588f9dfc802e7298a9a

Observation efae7ce4-c12c-4504-a690-d23a61462b1c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

LongCat-Image Technical Report Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.990675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:4f651a4b8a646ace02d375cef13cee1849f6249e1be32817281d89eec0b7eb1d

Observation 326de711-b8b6-4cf3-b850-786f6bf5f316 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LongCat-Image Technical Report Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.994233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:cc9e8b690e1eda25e4d44d8af2164f9b409d633ed6632d1a0bfb77cba6a233f8

Observation 0a7c3f98-7639-460a-8df8-06a00488590e · outbound

This paper cites Scalable Diffusion Models with Transformers.

LongCat-Image Technical Report Scalable Diffusion Models with Transformers

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:12.997672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:3d2a9e60dd611745a04b2b6b944454d46edefa8489689a070ce887dda4560902

Observation cd394c39-7167-421e-9a54-693af89a32d5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LongCat-Image Technical Report Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.000767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:f5b07a6396392b725fd78674a919172f73711d8c8148437be81c7204966eb11f

Observation 88b0bbfe-06b2-4378-986e-7363374a76c4 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

LongCat-Image Technical Report DanceGRPO: Unleashing GRPO on Visual Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.004401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:58b93b68044d07d9e356f0e147ec25e5652b3624ff93fa32c3bc5060ab56ff88

Observation fdbf2833-d313-4603-860e-ab0241fbcd07 · outbound

This paper cites Decoupled Weight Decay Regularization.

LongCat-Image Technical Report Decoupled Weight Decay Regularization

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.008387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:9e4ec741129cdc82a51f00bc85d9b2dd15ab5966e0ba6473c4f53ad72907718a

Observation 606ca2ea-c6fc-43fc-a781-1ced22ea0e5b · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

LongCat-Image Technical Report ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.012151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:bb2e9806acda3ee60e63d038aaed85a795c337d72dafcd8f387720b6e53ca01d

Observation 535c9a8a-8236-4335-94ce-44ad55be1482 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

LongCat-Image Technical Report WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.015483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:3abef5bcccf329fe9f87ea337e80fb2b6051c4f2f2350c5cfc10638adcf76d89

Observation 99dbe40c-b6ef-44b9-b666-b35b8ad4fbe2 · outbound

This paper cites Textcrafter: Accurately rendering multiple texts in complex visual scenes.

LongCat-Image Technical Report Textcrafter: Accurately rendering multiple texts in complex visual scenes

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.019639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:5fe7291127842aa4e302f5ab4e1ff3aae4fb077e2b342734e3f96d622f8818cf

Observation 8f0d98cf-46d7-4bad-8cfd-9b92acfb633b · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

LongCat-Image Technical Report Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.022951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:1c98c826161cfd52fe3169fb143785342e513f77e65ac157091852957b358212

Observation 5e24c089-f2e5-4c0a-9ad6-b7a4b9798dff · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

LongCat-Image Technical Report Emu3: Next-Token Prediction is All You Need

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.025983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:9637ea853f5b2e308b31223b22c28f8ce83299a26910cb2bba4a2f16ed90d559

Observation befe0f3a-7480-4cc3-aeec-63ec3f64f23f · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

LongCat-Image Technical Report Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.029742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:45d2b3af9fc3269c56ce5297664f106bf7cb43a648355280b086831d91ca419f

Observation c2697794-97c1-4072-94e0-0dc404fd7754 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

LongCat-Image Technical Report Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.033521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:826489b1d562fc98caf012f5f44416a4a432843ba0e83112f1807acd3d4da998

Observation 64b9d707-55ac-4cd8-ae5c-d22d46fc6f9c · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

LongCat-Image Technical Report HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:13:39.137651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:f020b8494083e29cfb26505cbe4e2a7a151fc58ce13d31a8331001ae5aa1c4d6

Observation 88a82248-2399-4ae7-a684-b0439fe82d57 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

LongCat-Image Technical Report Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.040150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:c5f0381e14ed6692d6f3b4df4263750e38d5ebe64cc6b4f2b56ceaf0bc343103

Observation 8b49922d-90a5-4d28-a7c5-cd4fdbd2b3af · outbound

This paper cites Transfer between Modalities with MetaQueries.

LongCat-Image Technical Report Transfer between Modalities with MetaQueries

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.044906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:4d7465ff2d1bb8e450fc3dab9c1758b9156cbc567e0974c8db5cf46328ba7201

Observation 1527e277-ac0f-4309-a7ab-9b8de4a040af · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

LongCat-Image Technical Report Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.052686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:9e9e1eb7331ad69a5b926467276ec9b75194f2eb39d5fb5bb62f4140fc3eb995

Observation fe42d59f-827a-41ae-b6d3-51a38373493c · outbound

This paper cites PaddleOCR 3.0 Technical Report.

LongCat-Image Technical Report PaddleOCR 3.0 Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.056034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:d4802c4ddf4a672ad5031e933e346647fa0d488c38b33c22c7445753be204bdc

Observation de38a0ad-99a0-4b45-b251-06a9f1714c97 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

LongCat-Image Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.062179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:e5924b35d15ea636746fc1f464b95e09d7bcd966b91639039dc61c8e5cfbf51d

Observation f8b6e2d7-fbe3-417f-b386-1e2656dc5bca · outbound

This paper cites GPT-4o System Card.

LongCat-Image Technical Report GPT-4o System Card

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.080196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:5ec04a3d4c37fd2db90d963bb0f898f1ececbca469184fd18d4a6fd5a830d4f1

Observation f941bc43-2b87-407e-adff-ac56ac07c517 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

LongCat-Image Technical Report Step1X-Edit: A Practical Framework for General Image Editing

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.083861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:4a8ba23ffb936e7403fd546909c82b0949eb7744f6262cb44ef9e6319566bddd

Observation 8b1a26fe-947c-4a29-9dfe-23776f869600 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

LongCat-Image Technical Report ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.087014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:9483d7d955a3859505a7bc3520b1bab52c1ee477dc39b71348a8815123413cd6

Observation b724bdf5-53e2-4771-980f-937035b540cc · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

LongCat-Image Technical Report UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.089941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:d4fbafc4fb2ce02d0514c838e0bb84866a6c892a6730b24de4f1513159408e32

Observation fe7def2a-3629-44e2-b76c-04308b5ee8fa · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

LongCat-Image Technical Report Emerging Properties in Unified Multimodal Pretraining

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.093121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:d6745b35d6d4b766386e86cdc0da2669a0585934eee4fedd48838b5c6263b17e

Observation c5ceeaa4-67a1-4e87-a85a-966990a45f27 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

LongCat-Image Technical Report In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:07:53.247084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:95e508c4684b8116259e9145dff9d8ba5139d357a20eab416cd30c9206cdf520

Pith citing papers

Observation 5bc87f17-5a4e-4882-8373-20ec8e90cca4 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LongCat-Image Technical Report

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:40:49.001525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:37:57.120779Z digest=sha256:e62c5a839c1f5ee7a4456ecc463d43a7b90157b64d2af5843b654c722d4992d4

Observation 2f636823-3802-4b33-bad4-320f77a7a7d2 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LongCat-Image Technical Report

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:10:13.506735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:167952bd62603364482e75d0e6c2826141b37410d7a70e36c71cb4c37bfa093e

Observation 14559017-f7d2-4222-9cb0-6ca298e98113 · inbound

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval cites this paper.

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval LongCat-Image Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T05:57:49.524269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:57:49.524269Z digest=sha256:00efa1107eca7b4353025b39625bff827748d3773c18c37e95e5efc27c86ef13

Observation 90ecf533-48e8-4a48-bf94-91fda535ab25 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing LongCat-Image Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:57.227130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:57.227130Z digest=sha256:eae5ba41cf2a689cb66535c86436ccfdca91d77a49bda8d764ac70ac847db539

Observation 4bcec20d-3e38-495e-9ea7-3d4b58c10b7b · inbound

Inference-time Trajectory Optimization for Structure-Preserving Manga Image Editing cites this paper.

Inference-time Trajectory Optimization for Structure-Preserving Manga Image Editing LongCat-Image Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T00:57:56.689235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:57:56.689235Z digest=sha256:f4904420d298a64b1c22c63475a9772c43a037fdf5537575b2686ab9d66ba4b9

Observation 00b2f9f7-1885-4fc2-8332-16bd102e6681 · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation LongCat-Image Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T21:18:10.087258Z digest=sha256:e4c796e64d75c7499d7aa420289fd69eaaf0e8b0a8d8add3a9fb2a5664504bc0

Observation 3c013827-5173-4366-b977-2bed29af7048 · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation LongCat-Image Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:35:25.116236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T06:31:13.232880Z digest=sha256:93a76a3ce9672f0ec3f0f157071339d3f907999f49b17508ae3cb51c657ec990

Observation d181c44f-f871-46bb-bd22-7d60c268b957 · inbound

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cites this paper.

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing LongCat-Image Technical Report

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:23:50.614589Z digest=sha256:82d817f7f39ac77d637b07feeffcbb1b4632194ae790f1ba01b61b9cc850775d

Observation 3d4f15aa-987c-4600-b718-ca68b24a34bc · inbound

FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding cites this paper.

FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding LongCat-Image Technical Report

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:09:27.385799Z digest=sha256:4c9b76268ddcb9585668876474757a2417a1c85da0f4045abccec47ffa1535e0

Observation e373609a-e2ec-4b06-aae9-af7f325bfa47 · inbound

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance cites this paper.

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance LongCat-Image Technical Report

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:15:23.578176Z digest=sha256:785011517307f809e1bac6741fba0f26a0be59fa7b4645f630130276a3d0344b

Observation 9acdaa3c-c9a8-4c7b-9325-639c6719a7ed · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model LongCat-Image Technical Report

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:9edd70c9a533bbc9b2a0572c36c1fcdb8289b6fef1190e82f5ebdaa81ed27284

Observation b86c029a-a496-4696-a796-f65e8e57c5b8 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LongCat-Image Technical Report

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:81d81bde01e929458197a39246e336d8fd00b4b71fa8f85fa5b09b09711a5b83

Observation 0454ea17-2581-470b-a4cf-6ed4919cf638 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LongCat-Image Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.192137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:b4b2ef3bf18ac8034563957dad26f5016a50f43e4cbfe3855c58e6d0f1d2edfc

Observation e6407dff-9705-4b86-afee-73f05f1833de · inbound

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing cites this paper.

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing LongCat-Image Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:44:23.162726Z digest=sha256:d2a00fdf536d8286d286092d8bd8ca769d2ea3fd717bab936abd44944de3a006

Observation 38adc94f-de51-42b7-a265-1707320edba9 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation LongCat-Image Technical Report

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:13500162a7ac1cf8d616d0ce24e596657ff9bf1e8ded0a3834819e8b7d23df07

Observation 8fb70b1c-f0f8-4f0f-9312-a5fdec7c5155 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation LongCat-Image Technical Report

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.745748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:1da2972c03706931738b6e317903775ba77f47f40853cf3a09935cffec2da2be

Observation aa3b25c3-3600-4fe8-a1f6-6b1426760956 · inbound

Open-Source Image Editing Models Are Zero-Shot Vision Learners cites this paper.

Open-Source Image Editing Models Are Zero-Shot Vision Learners LongCat-Image Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T16:44:53.286118Z digest=sha256:b3752b1de213c0e8668c5c36e498717d5b855e521ad93f9666e264acd586c9c0

Observation e69643e9-d56f-44ad-a19c-6c3c3eae8456 · inbound

DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models cites this paper.

DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models LongCat-Image Technical Report

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T13:54:00.141439Z digest=sha256:f8cac15a8a63e03835f04e13f594e2cc5634872bdfb916a7f6f5d937f317d448

Observation 9ee769e6-9cdd-4907-8751-873dbd41fb10 · inbound

Continuous-Time Distribution Matching for Few-Step Diffusion Distillation cites this paper.

Continuous-Time Distribution Matching for Few-Step Diffusion Distillation LongCat-Image Technical Report

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T13:28:42.083284Z digest=sha256:5c7751bc14a22f33ee14a5239178830c0e21bf58b6b64cd8de5e746558c12cd0

Observation ded175fd-62e5-46e6-b4de-cfdc1c52817d · inbound

Qwen-Image-2.0 Technical Report cites this paper.

Qwen-Image-2.0 Technical Report LongCat-Image Technical Report

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:21:18.312613Z digest=sha256:d371ba6bc85834d4707b98370e28de119092d85fdd85d54cb2d49507e0fada35

Observation 189392aa-2111-449d-a3ab-864cc4d77966 · inbound

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition cites this paper.

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition LongCat-Image Technical Report

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T07:46:50.540528Z digest=sha256:66fb3c221a2d354b4f580c02f72ab726cea01c12f163550e37bbf225f070a127

Observation 838d5c97-3ccb-4a72-aab0-da79130e38bc · inbound

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation cites this paper.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation LongCat-Image Technical Report

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:1100906ae0a1d5cccdd68f952237886d96b1e4a78249b87c8dd2cf0282083064

Observation 7651314b-648e-4ae9-9270-1635c76699fa · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation LongCat-Image Technical Report

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:940e40fb30bcdca5091439e2cb3315053b99116f9928d133701171eb918a1f4c

Observation 07e2ffac-9eff-4176-a52d-8a7b7afd1ed3 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation LongCat-Image Technical Report

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:d631db4361502d733fb2c972bdbe7fba0889e835a03cdc1cf2c038d00d3606d8

Observation f49d02a3-453d-4aca-8187-80fd158c7b87 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture LongCat-Image Technical Report

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:ba8e4fb2abdee3f349a84cb6ab6cf4bb84a70f84e52e02381d722de2f7ff179e

Observation 835d25a4-0e09-45a1-b499-f111c6765fde · inbound

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling cites this paper.

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling LongCat-Image Technical Report

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:22:20.464966Z digest=sha256:f6e240ce642e4eb2b124a8de8c46b18fc3abe7119e1304b329a41733b4a2e7ff

Observation 79025296-9fa3-4001-bca5-0f8d4a175530 · inbound

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation cites this paper.

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation LongCat-Image Technical Report

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:44:36.511647Z digest=sha256:3bd0086524583a7e071448f6ce6e11b5ab017a05f1cbbe0eb313e4d60dd586c1

Observation e02ce576-b94a-4e6b-94de-6e86c2a91a3b · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning LongCat-Image Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:37:39.780535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:02a1528189cc3fc6c5d20f43822300ab5168479336731e6918d8aedbb3224e0e

Observation 1e30b6f2-f8c7-4f24-9c66-166bed1bac8a · inbound

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning cites this paper.

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning LongCat-Image Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:12:37.557179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T15:09:47.549243Z digest=sha256:043243a30b2935194512f0e1b63321186a5f00e6d1aa70b2f05289f68e48fbe3

Observation 28635ad3-7819-4d41-a568-50d118776764 · inbound

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning cites this paper.

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning LongCat-Image Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:45:50.316447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T19:56:16.026242Z digest=sha256:572f3fd7eaf6501b10d24ab5b6a9e5d6b88c84a91adca4db4ed45fa8b54114cc

Observation c84aceb3-dec1-47a1-a36f-7825ee420759 · inbound

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models cites this paper.

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models LongCat-Image Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.609181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T21:18:03.005508Z digest=sha256:19d2204a0fe23fd59f7123e3c01349744af25028b7f798ca5d3bcae595146524

Observation 4aadfc4a-679b-46cc-80cd-91613848ca31 · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond LongCat-Image Technical Report

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:58:07.616421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:e1691de6d1a914a21c4fa267ca6af34ac6c3d472c4f0a1026aaa0c631fdf605e

Observation ad4f41a2-5759-4c78-95e3-59860989fcc4 · inbound

TextSculptor: Training and Benchmarking Scene Text Editing cites this paper.

TextSculptor: Training and Benchmarking Scene Text Editing LongCat-Image Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:19:39.364125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T05:16:43.756525Z digest=sha256:f23211ffbccdbacf1819a3296173aa1ff14aa934227aca7d1e94bff3860c7296

Observation b4d6b919-0d38-4796-b73d-596a3287f721 · inbound

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models cites this paper.

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models LongCat-Image Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:34:46.693481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:34:14.596976Z digest=sha256:d8cb86248bc962ffcf656f9577976b1c4eab50c68e63fc9d5cc1110deb1be8e6

Observation 678232cf-9163-4325-9864-dd1690ac849d · inbound

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation cites this paper.

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation LongCat-Image Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:31:22.978052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:29:05.360018Z digest=sha256:15ac2f46ccb37b5fff56c72eb8f73b4c8d55e59c1e28a11f953bffb4571b0246

Observation c2e572b7-5a8b-4559-8e61-dadaf06faba9 · inbound

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation cites this paper.

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation LongCat-Image Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:40:24.021806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:38:33.676797Z digest=sha256:7b9e67a134e64c7b78a24f792742994ffc7b4f82da5192b28f1af18a29616656

Observation cf4d177e-c982-44f5-8e03-0615c9a0dbd7 · inbound

ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement cites this paper.

ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement LongCat-Image Technical Report

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:00.855811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T23:03:05.701686Z digest=sha256:b384142fe365e5fdfd221dd80a3961e42fd35650115d9bc77a8b1d1b4f55ee20

Observation ff735590-a6e0-4e0c-924e-8ef88a703522 · inbound

PaintBench: Deterministic Evaluation of Precise Visual Editing cites this paper.

PaintBench: Deterministic Evaluation of Precise Visual Editing LongCat-Image Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:22:34.344342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T19:20:04.516202Z digest=sha256:8ef43912ac6f376a149b53b6dbff870c064073ca80098fe0b966b9bd4336c7ba

Observation d4cb7013-3f61-4b2c-bebd-7d4c41ae19b9 · inbound

TextWand: A Unified Framework for Scene Text Editing cites this paper.

TextWand: A Unified Framework for Scene Text Editing LongCat-Image Technical Report

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:36:57.272474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:56:46.514177Z digest=sha256:4e03d88bfcb8df00fed09bdc0f28c487c5a01b5865feeceb3206d817d0024c6c

Observation 95202534-ad41-4254-b62c-cff68b38a66d · inbound

Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment cites this paper.

Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment LongCat-Image Technical Report

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:06:55.947542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:30:45.673752Z digest=sha256:ac69bbf18a8714977f68a94e4588abff6e784bdec0df772f270083ebf91a807e

Observation 294d94fa-b3d2-49ff-9869-6e0b54006c27 · inbound

Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment cites this paper.

Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment LongCat-Image Technical Report

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:14:37.657074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T11:11:21.975534Z digest=sha256:224c37a140377e4e7632a12a20b73259e490296bdcb1f0c9cfc0d2f379b6bd22

Observation e510f435-3329-41ab-bc78-655d6f61a217 · inbound

InterleaveThinker: Reinforcing Agentic Interleaved Generation cites this paper.

InterleaveThinker: Reinforcing Agentic Interleaved Generation LongCat-Image Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:08:32.874817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:42:34.126336Z digest=sha256:b722370af89026d7eb836838d8f67fecb6937be52ec04ff942120c2e42315193

Observation cf3f6560-3849-4a3b-8fca-4f654ace8018 · inbound

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation cites this paper.

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation LongCat-Image Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:19:47.379911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:59:40.467458Z digest=sha256:92fbd4b3d047bedb73cee9451959161dd71578cd0c81895e62ed1b4ec2fccc6a

Observation 352d1d6b-bfa4-46c0-a239-11fb7f7bbd39 · inbound

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation cites this paper.

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation LongCat-Image Technical Report

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:09:57.575090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:50:22.839005Z digest=sha256:c44509f4763185411a80e005e0260531a2296aad79d4f8439fc185cd31565e66

Observation ae4fd954-3101-4971-9c0c-125f47c5121f · inbound

OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation cites this paper.

OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation LongCat-Image Technical Report

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T16:55:50.942261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T04:29:09.130025Z digest=sha256:008704616c9ffb4c5391df37a1f3a3c2a641668e9a42e434f67a8e755361df4d

Observation 86c0c891-9df9-4e99-b20c-ba012024ba53 · inbound

Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading cites this paper.

Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading LongCat-Image Technical Report

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T20:03:57.382354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T04:30:34.811205Z digest=sha256:19a0a244b75b14c555db1598856600be668868820d6f17eb676e87e80f6d159f

Observation ba7f2fb6-8123-4acb-84a8-afd2e07e67e0 · inbound

MirrorPPR: Exemplar-Based Portrait Photo Retouching cites this paper.

MirrorPPR: Exemplar-Based Portrait Photo Retouching LongCat-Image Technical Report

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:54:22.086899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:50:25.117845Z digest=sha256:6931c245c666ccf60a1fc8adb05e2cee603d9f18411a3d44491c050abb081f20

Observation 4eeb9ba2-7305-4c62-b722-c54a860c9905 · inbound

MindAU: EEG-Conditioned Facial Action Unit Editing via Dual-Stream Manifold Alignment cites this paper.

MindAU: EEG-Conditioned Facial Action Unit Editing via Dual-Stream Manifold Alignment LongCat-Image Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:27:04.662699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T15:24:08.345483Z digest=sha256:171fadbcefcc62b073987152b75f2c52558c9ada6c404b7d73d3eda751cf767c

Observation 3ac297df-67f2-4219-9b1d-e80704ae7a2b · inbound

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows cites this paper.

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows LongCat-Image Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:28:31.237417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T14:22:25.280140Z digest=sha256:002bd92d04930563b5cb99a1455645ae1a063bf1f0bcdf7647a0ba7c1df13599

Observation 741bf507-52ed-4f41-9e96-031c6dad2d31 · inbound

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection cites this paper.

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection LongCat-Image Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T01:35:43.979206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:35:43.979206Z digest=sha256:0bcc90c5b0d0f480863c584c69d41ef0abd26becf115b528ac19e1a4111e3219

Observation f3b5a891-e789-4d8e-aa27-55d9bce660f3 · inbound

DynEval: Holistic Evaluations of T2I Generative Models in the Wild cites this paper.

DynEval: Holistic Evaluations of T2I Generative Models in the Wild LongCat-Image Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T06:20:23.251121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:20:23.251121Z digest=sha256:8d22a15e7c1c85338ffb277b128ae51f1f95b6dca24b1e04978401d274784bcf

Observation 94a09113-b191-4060-95e9-c4a64c3caf9a · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget LongCat-Image Technical Report

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:03.536741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:03.536741Z digest=sha256:4e0a61dead341b2786f3d420cf3cf861438147c6ba764c849724473a6b32dfd7

Observation b7e2c99c-8ce9-4c0d-82f7-6fd3e7b8b8ea · inbound

SciForma: Structure-Faithful Generation of Scientific Diagrams cites this paper.

SciForma: Structure-Faithful Generation of Scientific Diagrams LongCat-Image Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T16:12:23.762958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:12:23.762958Z digest=sha256:266abd48ae7a1a105bf75e4d9276bbb4b25d06da6de0b8c80b018409cefd9d67

Observation 27429766-5990-48f7-bd3e-34c7de2fc427 · inbound

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric cites this paper.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric LongCat-Image Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.150238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.150238Z digest=sha256:99330a9bf68ff1b5f19aa71368bde0bb434d9a4d75ec9c70dd18fc896aec4e23

Observation 2f36cf89-305f-4a2a-8bea-5aff780c04d2 · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing LongCat-Image Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T13:38:58.078447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:38:58.078447Z digest=sha256:002ad850cbc7c9b5ff119585c69f47d72522379da8ba4a10cb3d7967cd27c027

Observation af7ab97f-e70a-446b-aa4c-1a7a7b0cf8e3 · inbound

Test-Time Curriculum for Open-Set AIGC Detection cites this paper.

Test-Time Curriculum for Open-Set AIGC Detection LongCat-Image Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:23.113783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:23.113783Z digest=sha256:1defc1cf32a9d9c3ceb6adb215833436372083a35543c71a68a1ceb648a780d9

Observation ba1db92a-2630-4efb-91ae-673c54486540 · inbound

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization cites this paper.

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization LongCat-Image Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T06:04:05.728155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:04:05.728155Z digest=sha256:bec41cd97939965c3a82c6b742a820d14dbe7eb5e8127ea67d3552ad21c09769

Observation 446143d6-f499-4551-b330-86af913451f2 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation LongCat-Image Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.542639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.542639Z digest=sha256:18ba7e9c2b6857a470661a079c8f583811fe246482851b8c34312768fe32497b

Observation 8691bc38-f3cb-4d90-83b1-b31ccbf23de2 · inbound

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation cites this paper.

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation LongCat-Image Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:38.987981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:38.987981Z digest=sha256:56126a9b7c31d30beab871abc13bc2e5a4f98e98c42f1ad3bd3ea4034e162cf9