Pith. sign in

Paper Citation Record · LEDGER

Jodi: Unification of Visual Generation and Understanding via Joint Modeling

As of 9 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 3 inbound Pith citation observations for arXiv:2505.19084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19084 v1

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:22.585226Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:38:39.360472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:44:19.355063Z

Reference resolution

100 of 121 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d01027f-d400-4688-8fad-0eb0ba28e6a8 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.023881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.023881Z digest=sha256:5538f5bf16ffb9e934bd7db75fb7c5e630c8399959cb07db45ed93173c567846

Observation fd6e124a-d999-445a-a6bc-8e4ceb437b5e · outbound

This paper cites Building normalizing flows with stochastic inter- polants.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Building normalizing flows with stochastic inter- polants

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.085409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.085409Z digest=sha256:4dbb4e763f7e1f40e3dade1285c30f29db3a4dd9c6445edf7b8eb744db9c9ac7

Observation a5191bef-b846-4461-8686-53a8bd4b77f4 · outbound

This paper cites SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:25:23.226183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:20.166543Z digest=sha256:88deee60f2060bec9eb0d46007cc33e7a06775447fb5ac10a991e2f5afdd6e6e

Observation 15bd8a39-1ea5-4cfb-934a-7ffc4a46fc79 · outbound

This paper cites Contour detection and hierarchical image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Contour detection and hierarchical image segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.269453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.269453Z digest=sha256:49ec3bf80e2e441d106b91c9e22941a6da5be6e7be27f2acce452050051bc480

Observation 0f7ea959-4290-480d-9d5c-248382f79c16 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.356553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.356553Z digest=sha256:4f4c9d55b4b981d1b1bbec9cb3b3a5dc670afcf429bafe0e658a66f240c8c0e1

Observation 502e35c8-1da1-43ee-98b4-f909b3f00b77 · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling One transformer fits all distributions in multi-modal diffusion at scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.438427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.438427Z digest=sha256:e1b67294beabf4fb9ee04724108bff981394d833e3a1bf73d8d54b3d71420cc1

Observation 79e3745c-1ebb-4f90-985a-310525f35c9b · outbound

This paper cites an unresolved cited work.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.520178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.520178Z digest=sha256:7fb25f689db526093c460f9e4b38ed19a4ee95153e8a05fcbcd739dc10363405

Observation 8e387625-7020-4d8a-a362-9e4e3bb53d7e · outbound

This paper cites an unresolved cited work.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.740440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.740440Z digest=sha256:bc83fbce4c3dd6be171a169a78de08c87c551b98d0becc37899f632c800cd07c

Observation d1a859c1-ac95-496e-b17f-eeb7c2fcfb13 · outbound

This paper cites Intrinsic image decomposition via ordinal shading.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Intrinsic image decomposition via ordinal shading

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.900857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.900857Z digest=sha256:1373e209c12d8f140dbca968741cb07b83174a4edebfdb45e2779a4f7d8c4ec2

Observation ebba4899-bf1a-4efb-b1db-2d78d61b6e55 · outbound

This paper cites Colorful diffuse intrinsic image decomposition in the wild.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Colorful diffuse intrinsic image decomposition in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.067772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.067772Z digest=sha256:88ec54edc115ada12b440d9ecf3279a3d63b7ec8a99889f638fcb32111d69eda

Observation 1b9a9ad4-5879-41ef-aeee-ce67b278bf62 · outbound

This paper cites End-to-end object detection with transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling End-to-end object detection with transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.233841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.233841Z digest=sha256:223a696cc9b2a546615e090443b96e4e06762a7b595ae66297d111f5b3425018

Observation a1fa9929-cb5f-498f-86c0-c5d7d6dec632 · outbound

This paper cites Artists as experts in visual cognition: An update.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Artists as experts in visual cognition: An update

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.361271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.361271Z digest=sha256:f51d8a6677e1609139b40a0c031dfcda798075a9c764233a6468dd17f95957cc

Observation d45b59d7-5999-4ec0-8189-892cbe76d218 · outbound

This paper cites Learning to generate line drawings that convey ge- ometry and semantics.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Learning to generate line drawings that convey ge- ometry and semantics

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.398707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.398707Z digest=sha256:5c196f415d68ad6302a24421040789f4f9b9f1f92ae97212c4b244a7fbaf4495

Observation e627f4f8-64db-4e6e-9475-432074c64545 · outbound

This paper cites Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.463272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.463272Z digest=sha256:4ce934ecac475fc03304c403c34de4e25fa9a1d74ea4bafd76cb7eb91db97d0b

Observation d218a445-6150-48c6-963b-e6f49d8505a7 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Deep compression autoencoder for efficient high-resolution diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.609686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.609686Z digest=sha256:bd0b48a6928984c51dca92265072d00231b3cdee2cb63335267d59c854bc7f94

Observation 456e047a-67e1-4267-b9d9-fd4357d45170 · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Anydoor: Zero-shot object-level image customization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.659977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.659977Z digest=sha256:4a54ceb88c685b0a1734d5f4a88a1eb8621d6ad43543d2fbe9573d327410883d

Observation ab528964-950d-4149-ab96-c82f5e05e0cd · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.664689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.664689Z digest=sha256:15c6c12ab4598314cf8ffdf7042680014b670c076acfca59c4f3b05b45a71846

Observation 2d08a53e-606c-4d05-b200-576eb63a0664 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.733104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.733104Z digest=sha256:3e1e71086884cc3f2737668f5a585890ba321c46fd8ae099af3ea5045353dd10

Observation 3004283c-7d1e-4f43-ad3c-f69b81981ff7 · outbound

This paper cites Idadapter: Learning mixed features for tuning-free personalization of text-to-image models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Idadapter: Learning mixed features for tuning-free personalization of text-to-image models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.789469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.789469Z digest=sha256:46fd0c3266ce1bfb692af98958faf7e96a5dca2245c721745913cc0f8a3b4e5a

Observation c334c49e-2a5d-4c1f-ab37-8a2fecaaac47 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.869510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.869510Z digest=sha256:38df687e556539d306ff869ce8fc9488da4702f0724c514bf8ce18d7988615d8

Observation 742df795-4d6a-46ce-9df0-3bbc4911245e · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.956584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.956584Z digest=sha256:3a47aeefaef06a9fbf5a3a55907ac6c1503778feb14b38257ce4ad018829e195

Observation c9a7aaa5-c1c9-45d0-bccc-974431354446 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Imagenet: A large-scale hierarchical image database

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.030338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.030338Z digest=sha256:2ae0b19ad436f3d306ba4aa211bfe0eadc34207c8aaf4ffd539f9b52618226f9

Observation 19e0a22e-b91e-4bc0-8275-974ff4701e59 · outbound

This paper cites NICE: Non-linear Independent Components Estimation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling NICE: Non-linear Independent Components Estimation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.092599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.092599Z digest=sha256:5e6151c3918b029395a00da3e04a9a5e2a647a647aaa6c329415bf0029ecc38a

Observation e1d25dda-9a3f-4f47-bbb3-9cb702117117 · outbound

This paper cites Evaluative and generative modes of thought during the creative process.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Evaluative and generative modes of thought during the creative process

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.177856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.177856Z digest=sha256:e3af81ecb2cf9b97ae2e103f2a67ad9a132dab7c50f815e1556e812040345cde

Observation 451cda10-37cc-4624-ab4b-2565953a710a · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Scaling rectified flow transformers for high-resolution image synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.229147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.229147Z digest=sha256:86f009b2ab35ac40044f78c7cf175334a07303088a6cb0c7f2d6197ab4f936b3

Observation c300ea3e-c3eb-4861-946e-6dbe3e3c14ac · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Taming transformers for high-resolution image synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.233468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.233468Z digest=sha256:3256e6bd29695220f430c2f506a3df197803e32fcc70c7a32bec265ab2402195

Observation 919ea3c9-db4c-4b48-ab61-07804eda1506 · outbound

This paper cites The surprisingly powerful influence of drawing on memory.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling The surprisingly powerful influence of drawing on memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.238046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.238046Z digest=sha256:d8ed1407a971b1decdd8ef8016dec9a02e97551b2ed8dd691948b1ba97854e12

Observation d2a5110c-c8d9-4efa-bc82-2ef113613f0a · outbound

This paper cites UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:25:23.130205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.241809Z digest=sha256:ed8450b9d8c6ff9c64c5204ad0ecc16365b5d4ab043136d0d8c0ddab206f383c

Observation f2637597-011e-4ff8-a947-2307739055e9 · outbound

This paper cites Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.247783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.247783Z digest=sha256:dd4219797fbe19db72a749138314471e8bc8fc807f7d163b5afcd2d3ba192a6d

Observation 4cfd5468-25f0-4cf3-9061-b3120997a6ab · outbound

This paper cites pexels-portrait.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling pexels-portrait

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.251756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.251756Z digest=sha256:928f85f217c5609cd79001dcb96108ba55854d9a0bffc0d7d8453c34027a8ea3

Observation b9425971-ce73-4695-b7c9-2666ab61252d · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.256655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.256655Z digest=sha256:fdcaccaab0c4ccc5340dae1a668fe01a2963fe24851e7a049f91792d3332ca09

Observation 82a39014-5b84-4f2a-b501-b3ef600bbcc1 · outbound

This paper cites Generative adversarial nets.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Generative adversarial nets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.261526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.261526Z digest=sha256:2d52f5204f1741c1c1a176be8185d80e04422d8c06175607a0a79928b6314045

Observation b2cdb417-a08b-4101-a158-e3e32d41f347 · outbound

This paper cites Pulid: Pure and lightning id customization via contrastive alignment.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pulid: Pure and lightning id customization via contrastive alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.266396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.266396Z digest=sha256:3c21cd33e8e2d985d02ced85e0e4a5667c69a0210402f78d8ddc6d36b080bcbf

Observation 26c8e1b9-7923-4a9b-a678-10a051576392 · outbound

This paper cites Svdiff: Compact parameter space for diffusion fine-tuning.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Svdiff: Compact parameter space for diffusion fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.270883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.270883Z digest=sha256:d16f400293a41cc4bc0d313025c59589c698d05290c7e68a57e81bf00f279d88

Observation cfdaebcf-74cb-4790-8fa6-2bfa9d0d0e4b · outbound

This paper cites Transformer language models without positional encodings still learn positional information.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Transformer language models without positional encodings still learn positional information

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.276191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.276191Z digest=sha256:9f19db5b3b5388f007be344152e815ef4a3cc019a30dfdc9f8db91a8451172ae

Observation cc1ec785-7b50-4e25-80c2-e7ce58fe012f · outbound

This paper cites Lotus: Diffusion-based visual foundation model for high-quality dense prediction.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Lotus: Diffusion-based visual foundation model for high-quality dense prediction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.280980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.280980Z digest=sha256:e05a2fbda541ee344d7f9ef321a5a787073089f35eaf3960192bb9525e9ebce1

Observation d36af7db-ff5e-4872-b15b-6a7a128f33d8 · outbound

This paper cites Mask r-cnn.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Mask r-cnn

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.285397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.285397Z digest=sha256:cb45ae6373706cadf7408e4e7582ebcaa17f7f319fb78e5053c2131bd34b8e03

Observation dc5859f8-72e4-4d0f-92e7-42cf56f6af5f · outbound

This paper cites Deep residual learning for image recognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Deep residual learning for image recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.289875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.289875Z digest=sha256:785f543a4fd3bb43e3e9dd8ae8b519707507e5814fdb7cb8b67755a2c09f2d47

Observation 7e05d896-e9fe-4625-9376-779459bde742 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.294337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.294337Z digest=sha256:8db10e186c64a24e7a0355c665596efc3cf6c0d7dd95c06236e8b048ae867ba8

Observation 0a924934-76a8-44cb-a44c-61e7dae76d92 · outbound

This paper cites Denoising diffusion probabilistic models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Denoising diffusion probabilistic models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.299329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.299329Z digest=sha256:3d7105419c95480d8db085b4c9ad7411cb94b99be07c2467896aeea41611289c

Observation 588a842f-3af9-44e0-a282-3f7e67eeb00e · outbound

This paper cites Classifier-Free Diffusion Guidance.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Classifier-Free Diffusion Guidance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.303298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.303298Z digest=sha256:29978557b68cfcdc5294dde44000af04e39c999192ef29ed203e97cd9eb30aec

Observation b920f5c5-b7ba-4a81-a2df-94ea4fcba894 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Oneformer: One transformer to rule universal image segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.307733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.307733Z digest=sha256:c66ca312cf2c65a95d8546690cfc5351d21ec15fe23309080d75349eed98fe4b

Observation 5c156863-3be1-423c-b09a-805b9a3a4ea6 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling A style-based generator architecture for generative adversarial networks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.312506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.312506Z digest=sha256:2167bac7997f9f99748b32e516b3f87472cf2b32c8d1a255c9d15d2e324a8400

Observation 2a57b08c-6051-4221-922b-f2714d0086aa · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.316794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.316794Z digest=sha256:4940e73756d0aa99eef24756e4012bbedde779a6cb39fafbb385aeaf9804d41d

Observation 00bf20dc-1533-48c0-a6f9-3f3d7b491d0f · outbound

This paper cites The impact of positional encoding on length generalization in transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling The impact of positional encoding on length generalization in transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.321110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.321110Z digest=sha256:7ca1e023a519c7d92fbbfa40b9e2ffbc11aad0e886fe6e4aabcdbfc969d2658b

Observation afd4cb01-edd3-4521-a8dc-75d8958e415e · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Repurposing diffusion-based image generators for monocular depth estimation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.325277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.325277Z digest=sha256:4f902d20a9ffec22e28fd08150992f63c3f8d8f98ecc105803e10203c4c9501b

Observation e274ca9c-b875-43ec-92fd-4143dcfacc17 · outbound

This paper cites Auto-Encoding Variational Bayes.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Auto-Encoding Variational Bayes

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.329422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.329422Z digest=sha256:e03353c91a3df0651e78a4514fcc3f1507c8df127dd948aa4120aa11bef7ce50

Observation 811f385e-1963-4c3e-bc59-606dd577116e · outbound

This paper cites Segment anything.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Segment anything

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.333436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.333436Z digest=sha256:56a84dc097e8f2a329bf99bebe499403e805347fded083c00482b7973d0dac97

Observation b6ad9372-3893-4659-8528-2b773675776b · outbound

This paper cites Evaluation of cnn-based single- image depth estimation methods.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Evaluation of cnn-based single- image depth estimation methods

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.337682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.337682Z digest=sha256:6dc4353801ff623e07365b3f013ef221f90813fa6be91a570f0258bd44b23d1f

Observation 238b022c-551c-4b10-ac53-bdc8f0b4133f · outbound

This paper cites Intrinsic image diffusion for indoor single- view material estimation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Intrinsic image diffusion for indoor single- view material estimation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.342567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.342567Z digest=sha256:0ca7d20086cd74278c19ed32c15985bad572f7ecb09102d3aa9f9f54d26fc9c8

Observation c7917e06-628a-4d5b-b870-84caef2c7d8e · outbound

This paper cites Artists as experts in visual cognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Artists as experts in visual cognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.995509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.347021Z digest=sha256:1548b3fcdfa72882447195da8126c1b224a0eeefdbd7c6860f5c7c3569509c07

Observation d6f11f3c-cfe9-4b85-ae82-90d39917510c · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Imagenet classification with deep convolutional neural networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.351696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.351696Z digest=sha256:9f6934b49463f2be1156ac92a2bfe7227a77b3f41139e621f3eeb3cd45eb36ae

Observation a3265254-8ec6-4d35-8701-b11f459d5907 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Multi-concept customization of text-to-image diffusion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.970690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.355703Z digest=sha256:a78562c102f891a05b5ac58edf44b274a7b7f9dd4cc7596232c678416bde89a5

Observation b1704c97-8bcd-42c0-9e37-2153ce3b547e · outbound

This paper cites One Diffusion to Generate Them All.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling One Diffusion to Generate Them All

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.360463Z digest=sha256:9eac2a812eb6c887914f858c4a5fb716e8d0fa3146197317b794cd2ccd48f874

Observation aafa071c-ad5f-438b-a2c7-8344e83d445e · outbound

This paper cites Gradient-based learning applied to document recognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Gradient-based learning applied to document recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.364853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.364853Z digest=sha256:8030b2336d82fc3a502b74bce60139568a438171b19962e3b8517a429c923360

Observation d1d375e9-e3f1-4e95-a7d8-5a16e3ce944e · outbound

This paper cites Exploiting diffusion prior for generalizable dense prediction.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Exploiting diffusion prior for generalizable dense prediction

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.946089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.368903Z digest=sha256:2eb40e97147c183a50cd6be04f86e564d6e6f65c979c285acbf29985424806ca

Observation 507d85d1-50d0-4329-8c30-dc4717ec4569 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.373168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.373168Z digest=sha256:cf66ef6e264f1785828c48a2e95cb7338d9586918f88f7fddb4a002905ae1b63

Observation 0e19f149-40c9-4c0f-a95f-f08d3513ec3d · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.378612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.378612Z digest=sha256:4e5d35fd1f12624d701177688c9edad58f2708769cbcd3e97a1bec558368eb52

Observation e493cd5d-fe29-4316-8fac-c760f4498989 · outbound

This paper cites Uniformer: Unified transformer for efficient spatial-temporal representation learning.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Uniformer: Unified transformer for efficient spatial-temporal representation learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.919950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.382933Z digest=sha256:01c2c4e1b21ac3132702144c388369ff525f72b8eabc0f8ddc7323569b16c195

Observation b66de16d-e7a1-485d-8e67-33dcedf2e603 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Gligen: Open-set grounded text-to-image generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.386866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.386866Z digest=sha256:b5c2c1c873ab5bf1bfa4518188a5c5e8f10a513c96bbe2f9086ea45a2e940c50

Observation 46c467f9-78f9-418d-bd47-fdb1a38194df · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Photomaker: Customizing realistic human photos via stacked id embedding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.390909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.390909Z digest=sha256:5a94bba55dabeae45ea135ec12589209f77156fca5bd86563712cbe2c3fd91ec

Observation ba11db8f-ae42-4811-9e68-11470bdeffb9 · outbound

This paper cites Dual Diffusion for Unified Image Generation and Understanding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Dual Diffusion for Unified Image Generation and Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.394934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.394934Z digest=sha256:ba17cb3c82df4da3bdb130b951a7f1f681a1af7dcaeca3aac3ebd17a143afb26

Observation 3caba455-76f1-4bac-ae07-9f4787975fa9 · outbound

This paper cites Pixwizard: Versatile image-to-image visual assistant with open-language instructions.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixwizard: Versatile image-to-image visual assistant with open-language instructions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.880370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.399786Z digest=sha256:4ae202f0dd5dfb6fb0001484f6e252807b9c2e89be91af8d51e32c262a576b08

Observation 27f252d0-8c15-4085-8ccb-812893f5a6b6 · outbound

This paper cites an unresolved cited work.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.406222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.406222Z digest=sha256:c5956cae5dcbd25c21e6704df267dc37f615710ac93d4dcf19f5df2486c6ee8a

Observation 5de40839-52c6-484e-92d3-ea71107275ed · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.410437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.410437Z digest=sha256:41a85a75872f8e00e65327fdfb032dbd3f62f6c13edea262e5298bda62a48d88

Observation 9f9e7c7d-a976-4d9d-b215-9c9a3bdcf9d6 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.415198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.415198Z digest=sha256:29e081f6733d286bba87a3a91fc92ce1ab66fc8321d9491f374cdf4643a2f8a2

Observation 91527882-fa98-463f-aa4b-9ab807cdd40c · outbound

This paper cites Came: Confidence- guided adaptive memory efficient optimization.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Came: Confidence- guided adaptive memory efficient optimization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.839289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.419555Z digest=sha256:3fa04f176b907fba5976b3d5ade1a9ec0b51f7d8a6446891d9adc5fad12b438b

Observation 6950d879-870c-4d4c-88fe-fa61443f0703 · outbound

This paper cites T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.823402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.424091Z digest=sha256:854a7cbee4808ccf4199be6d4a9b4e7099324b208d35f8111eb5999f3a33c512

Observation 1aec3c3b-0be3-4a5a-8f91-3c1392630109 · outbound

This paper cites Novelai improvements on stable diffusion, 2022.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Novelai improvements on stable diffusion, 2022

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.806790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.428891Z digest=sha256:e7bf02697669a67f85d715e4bd558ecd9280ffc9f06cb95b9476f56da940a697

Observation 8253015f-aebc-4a08-ab85-6d855b7596ed · outbound

This paper cites pexels-photos-janpf.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling pexels-photos-janpf

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.790583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.433006Z digest=sha256:4c3463ec6c9e11e0e5284563cad1ef4809cf523fff1b2e08c32de1ae0cc1fd01

Observation 56a94a99-f0eb-4314-905d-95b0d1394218 · outbound

This paper cites Scalable diffusion models with transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Scalable diffusion models with transformers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.437358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.437358Z digest=sha256:47f3736a9293c1ef4460f08015c9b9a73480b7ba0d56d85b7f86d0accb1c9618

Observation 6a4b3458-4594-4c27-9982-d82d246bb741 · outbound

This paper cites Ld-znet: A latent diffusion approach for text-based image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Ld-znet: A latent diffusion approach for text-based image segmentation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.764838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.441752Z digest=sha256:e6701d745e52ef8804270ee1cb295fcbc2fd711d73294524fe3b29bb35c4867e

Observation 536825bf-274c-4163-9677-33421549d7ab · outbound

This paper cites Unicontrol: A unified diffusion model for controllable visual generation in the wild.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unicontrol: A unified diffusion model for controllable visual generation in the wild

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.749161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.445661Z digest=sha256:21d41dcf60bff106ac1527f8c8a4db26aecd7943fd4bee939c587eebb82ab59d

Observation d51d9076-fb7b-4258-ab6c-d6be16c01cc8 · outbound

This paper cites Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.450236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.450236Z digest=sha256:4200e3e838bcedb44daab1090ec1c9fdafc741114eb6a643f3355a34d39fda22

Observation 705116ba-f314-4f5a-ad56-5f5e5e3dcaae · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling High-resolution image synthesis with latent diffusion models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.725331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.454569Z digest=sha256:209f885e6b96e9a4b4b1a8903ca17bc9a7880e99988352d11765c977e93ee4c6

Observation ad4a8edb-860c-40a0-8a2e-2add47940beb · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling U-net: Convolutional networks for biomedical image segmentation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.459677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.459677Z digest=sha256:b898af92c962b389493d92dcf69a386487f764c94f5200879bf845b84511d53b

Observation 24b5f3a9-af8b-492a-9e43-d8bbef430a82 · outbound

This paper cites Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.464421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.464421Z digest=sha256:81e547c3b20858c72e018e0a031b137627198b7231956e561327788b25d5f2cb

Observation a3a4b2a4-2ab8-4b2d-b88c-b7220352a623 · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Photorealistic text-to- image diffusion models with deep language understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.469175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.469175Z digest=sha256:24b2099e5b6b373cabbac7886b38aee242fd753f3b2397707152cb85289a3365

Observation 2a99b0ee-69d2-4ef3-a72d-e31ed09614f8 · outbound

This paper cites Ziplora: Any subject in any style by effectively merging loras.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Ziplora: Any subject in any style by effectively merging loras

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.474251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.474251Z digest=sha256:bc1f4b1d9dde6c2986e80c680a4c169ed85f3e0bc1d985680902af8af0c9c576

Observation 1337a926-c7a1-421f-9c38-17ae3e0d4cbb · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Indoor segmentation and support inference from rgbd images

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.478989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.478989Z digest=sha256:e65ba26ba67ed39ce526ed4a1d4645377f7005d48b68b82e9a1fb29d65a9c32f

Observation 33529430-2f5e-499e-9e35-4a05e2ef5d84 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.483167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.483167Z digest=sha256:9a0e294caa3b0026fd89433bb3165da68df1bf7f52f1453f259b16e3e0be9f55

Observation 309b4ac6-122a-4401-8f2d-fbbc8b0a387f · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Deep unsupervised learning using nonequilibrium thermodynamics

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.487987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.487987Z digest=sha256:b6c6f03bae93fd1302c1bdb5dd1f919e3e7104e60d76d1d911c165b65ad2b920

Observation 9223c140-660f-4f3c-bcc3-1aefd4a45fff · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Generative modeling by estimating gradients of the data distribution

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.492170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.492170Z digest=sha256:84e65d492419170c13580854d089c689b951c7282c0cdff64ad27fa4850291e6

Observation 4fe99b11-47c5-45fe-b27f-3e25421545bc · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Score-based generative modeling through stochastic differential equations

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.497026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.497026Z digest=sha256:824c034328fb889afc5409672b9dad1a256f1194d8788038662243b921ae2733

Observation d110e3d8-a2b7-49a3-93a2-764e51034878 · outbound

This paper cites Pixel difference networks for efficient edge detection.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixel difference networks for efficient edge detection

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.625294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.501389Z digest=sha256:b014d174957a7f7bf079476236ea06fffb8de889b2331065583d3e344ccf320c

Observation 9ad2a344-228e-47eb-8db8-265851a07630 · outbound

This paper cites Going deeper with convolutions.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Going deeper with convolutions

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.506402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.506402Z digest=sha256:d1b501413a4dcf208757250274874cc3c764c7838d1ae55168f33c7650582871

Observation 4d27d7ac-3902-4056-abd8-684fa6c21255 · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.511560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.511560Z digest=sha256:c39bcae500406f35fb6f05df33ae67bd1a4b6f56296de050331eba9685eab3c3

Observation 579ed44c-06ed-4cfa-8bf3-9eb1ed1e39e9 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.516710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.516710Z digest=sha256:45fc7f9a500812741cb3eeeaf7c2d3a17072b495d09c357543882cd907282ef4

Observation 308e4df1-654c-4046-91a0-dc895d6fc5e7 · outbound

This paper cites Pixel recurrent neural networks.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixel recurrent neural networks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.521960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.521960Z digest=sha256:50b2a5dc2a47599f2444bd8f70ac4b0eeafd5a08ec70a82969fc3b1607ae6797

Observation 704e821b-f21d-4751-ac9a-61d05220b4e9 · outbound

This paper cites DIODE: A Dense Indoor and Outdoor DEpth Dataset.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling DIODE: A Dense Indoor and Outdoor DEpth Dataset

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.528715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.528715Z digest=sha256:d90354d27640c7f7daebd5f85496e87b78e30b780d08a21e596fb25d442900fa

Observation b2f1a476-2042-4334-832d-01582ae434cc · outbound

This paper cites Attention is all you need.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Attention is all you need

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.534394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.534394Z digest=sha256:c87b40b91a881bb27477e853449a970e87b7b684a5314dda3e4e3427f508d464

Observation a9fa7bfd-06dc-4760-afd8-4232533b08e8 · outbound

This paper cites MMGen: Unified Multi-modal Image Generation and Understanding in One Go.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling MMGen: Unified Multi-modal Image Generation and Understanding in One Go

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.539339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.539339Z digest=sha256:4a609e78f1f60d0ea1b68f5ccef873e21548e8b21ca37c64a5127b7597540137

Observation 8b67ff78-b67f-4e1f-b116-53a3f2a43fe5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.545579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.545579Z digest=sha256:506be112168a748df653ba4eb7763289c013ff910d1d15bf172514273162e13d

Observation 6dabafcd-fe87-4a07-88d0-f7ce63c883d1 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.551869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.551869Z digest=sha256:6ce3f656d6bafe3be1a54c4e71c4546df35f5e1c87c0a1e69c7c5a7631234802

Observation 28729510-947b-4d84-9e2f-ba67af8485af · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Emu3: Next-Token Prediction is All You Need

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.556883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.556883Z digest=sha256:c55e37be1f0df2ead721b9b41fbc9f5ec75f0eb079e3aaa2576451564ce61e75

Observation eed5e9f7-4f3d-49e0-890f-3445f1aa2a37 · outbound

This paper cites VILA-u: a unified foundation model integrating visual understanding and generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling VILA-u: a unified foundation model integrating visual understanding and generation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.577141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.562233Z digest=sha256:8e76973fad34a306959b6858882d7eb3d67671c66fd01973f4461e09f43e4a4f

Observation 2ef3d780-38c4-4624-995f-ee8b35e92789 · outbound

This paper cites Infinite-id: Identity-preserved person- alization via id-semantics decoupling paradigm.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Infinite-id: Identity-preserved person- alization via id-semantics decoupling paradigm

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.559406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.568323Z digest=sha256:72e3c1c6ea30ec839bf75c893bb14dd21b49dcb97d8c697716ea2d1aa47b41cf

Observation 69893bf6-8d16-49ce-b7c5-8e94aad087d7 · outbound

This paper cites OmniGen: Unified Image Generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling OmniGen: Unified Image Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.573243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.573243Z digest=sha256:b4daf68901627872f46d9d07adcaf57741748eb689fa3ca5d7ad384a33bb3251

Observation 2978706d-23ba-4e0d-b4dd-3c23fee37374 · outbound

This paper cites SANA: Efficient high-resolution text-to-image synthesis with linear diffusion transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling SANA: Efficient high-resolution text-to-image synthesis with linear diffusion transformers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.579652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.579652Z digest=sha256:4336903b070e15c464201aff1c647a29a1cac851b911f940c34540b40f7ad222

Observation d8910e70-353f-4bd1-b693-250a93b310ce · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Show-o: One single transformer to unify multimodal understanding and generation

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.533825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:22.585226Z digest=sha256:59ed6752f70dc25e54d611c44944cbce9e5ba28f654f61b11395731e40f0f5a4

Pith citing papers

Observation bd0bfbbc-735c-4fab-8c87-1ea894e2b05c · inbound

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors cites this paper.

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:07.834476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:05:21.723724Z digest=sha256:dec685ad24557e2b9d734df5f513ba8b03705cc907ad250b9c82be40d5476c14

Observation 31191285-971e-4c0e-862f-aad9e6c2fdd3 · inbound

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study cites this paper.

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:25.664198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:53:06.362430Z digest=sha256:029bbdd5950bff1027bd77f75e2b3549238a9ff35a4dddb92f6e13a22744dac8

Observation fbb715a8-beff-41b8-b7c4-03b22b9d9b89 · inbound

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception cites this paper.

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.357185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:38:39.360472Z digest=sha256:c36621aa61bbff7c690cda949d18071eeda4ada374abcdc6c8afc9b53b3a2c41