Pith. sign in

Paper Citation Record · LEDGER

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2507.02321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02321 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:35:54.220222Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:47:22.111357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:27.320514Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36a3d54b-6f99-4e8c-b0f2-0c7b77895d91 · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:58.842070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:49.477443Z digest=sha256:884bd532335fb96be635c815b9740d2a1ccabbed6349dc25403427a419610e79

Observation 6b4632af-188b-4b6f-b8d6-efe4715ade8f · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Adding conditional control to text-to-image diffusion models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:49.524548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:49.524548Z digest=sha256:8f1b0b2344277556ad5f024446de48cf4061c76779dc2ddae9337cbe93b5622b

Observation e0a103da-e02b-400c-ba1c-708af90c3546 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:49.586081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:49.586081Z digest=sha256:a92cd476e85919096d4a1e5ecbe093497e5b2b78829f59df9a144d1bc7208f61

Observation 4c81f357-ae0e-49e4-a19f-be14bea1fbeb · outbound

This paper cites ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback: Project Page: liming-ai.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback: Project Page: liming-ai

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:58.619017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:49.652254Z digest=sha256:72ef0ef6c000a6ef80bf1c2f8b2d0f65c3283a93cb81ce646082bfd320825520

Observation 343056b2-3b41-451d-8cbe-ae0710410cc6 · outbound

This paper cites Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:35:54.913613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:49.706128Z digest=sha256:0c89f351cfabec389f18dda501a797e1fd94c1a73dcf59cf1206f302b3bf04e0

Observation af2b2fa6-d007-47b9-a69e-66ffdac07626 · outbound

This paper cites Denoising diffusion probabilistic models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Denoising diffusion probabilistic models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:49.790142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:49.790142Z digest=sha256:c0825974ec354c70a40d2de56a4b9e382fd82abc84e14d1d6474e2130d4e6c53

Observation 5971fe87-45e3-4af1-bc62-fa56c7f15b22 · outbound

This paper cites Denoising Diffusion Implicit Models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Denoising Diffusion Implicit Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:49.863663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:49.863663Z digest=sha256:343e69c50eac0c1ad443b4683d0a81666eb2a50b47624a97d8a2c077ffb5fa73

Observation 7c13a7d1-7fab-4a90-9b76-7ca08dfe365e · outbound

This paper cites Diffusion models beat gans on image synthesis.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Diffusion models beat gans on image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:49.935106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:49.935106Z digest=sha256:5390568d0be2d6031d6ab6b467bb9ea07423a523a0f758d8ff2e1ab5b5bf2d24

Observation fa84931d-d19a-432b-990e-8b9f831d1107 · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:58.409609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:50.076470Z digest=sha256:cf37c4018eb2c07983a642a6c1f9396f04cfa520497df5a6967d2953727c825a

Observation 5a81e00c-8471-42b8-b936-7ef4c26b2f01 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:50.202595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:50.202595Z digest=sha256:97dcb11c5dcaa4871978f454de9ff3b81d49fc49fface74813df8f6df55d613e

Observation b918a55e-1b70-4f23-acc5-b278d5d8856e · outbound

This paper cites Zero-shot text-to-image generation.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Zero-shot text-to-image generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:58.207271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:50.305528Z digest=sha256:21e5ee4c6c853e2cba1dc43b67d0ca52f931fdf1e992ae88b833baaeb7c6009c

Observation 6fa6221d-33d1-4be7-8655-ef56f987d737 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback High- resolution image synthesis with latent diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:50.417376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:50.417376Z digest=sha256:efb0c39a962761fe4b85c53839ef72a0071ed6d8bfba7c17aa4710652605101d

Observation 54958643-4533-43aa-aa33-179d7f9fd828 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Photorealistic text-to-image diffusion models with deep language understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:58.002866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:50.552056Z digest=sha256:b48394e8051c408b46e7d5d04db9430b21fd2f9fec2b1485e537f2fa45f259f7

Observation b08804e5-8ff5-4084-b159-68d3b0b1c927 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:50.644902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:50.644902Z digest=sha256:361190c97f2db803cbcf326e796abac8a8c323d3c8d46a08da5205b8f76ef6e0

Observation 18b9f45e-8175-4f4a-b3e3-4a13efa3c54d · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:50.758018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:50.758018Z digest=sha256:a06116a9fc52ad82cdcbb653591deb1562428a540127c68f6f2fa2e89be0bcd7

Observation 51876550-2700-4d03-8a44-ba567216a1d6 · outbound

This paper cites Cocktail: Mixing multi-modality control for text-conditional image generation.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Cocktail: Mixing multi-modality control for text-conditional image generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:57.831530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:50.877805Z digest=sha256:a0b60e6113ca67c1df80ac68b1b9b6b794e10a34fce52eb564d843bc7f958f1f

Observation 493b943d-5f7d-4e00-b84f-b779852fd43d · outbound

This paper cites Composer: Creative and Controllable Image Synthesis with Composable Conditions.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Composer: Creative and Controllable Image Synthesis with Composable Conditions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:50.956752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:50.956752Z digest=sha256:63fa5ea1c4c1b3f746e2b88c271d79735c7158fd02bfcb27d05711159df3a9e0

Observation 6817ac18-a19f-4c93-aef8-dabb0e74b87d · outbound

This paper cites ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:57.663335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:51.053687Z digest=sha256:3a59e4150baf65099f4e84d93f7157a3fc65a80f6dcdc21be97072a223bec6d3

Observation d29f8acf-1c65-4a8b-94c2-3eba77a736cc · outbound

This paper cites Relactrl: Relevance-guided efficient control for diffusion transformers.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Relactrl: Relevance-guided efficient control for diffusion transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:51.164157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:51.164157Z digest=sha256:21f04b85a6e231d67bbbc0c86217d85383d68307f44c3ee432649b44051ce86a

Observation 0254e3da-9d39-4408-876e-b64a853291e0 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Uni-controlnet: All-in-one control to text-to-image diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:57.496019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:51.315400Z digest=sha256:76a2b67814267a94d7361a6f1d87cf030d42db8b01ae883bda183d118411bfad

Observation 244f393a-58ff-4d2c-8f13-d54703c61389 · outbound

This paper cites UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:51.439611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:51.439611Z digest=sha256:40d773306c3f301eee96a65a89d1cce002ecddac302f57fbfe31bb32c3dfdec1

Observation 00051692-5970-4a13-aaa0-090f4d048241 · outbound

This paper cites Beyond Surface Statistics: Scene Representations in a Latent Diffusion Model.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Beyond Surface Statistics: Scene Representations in a Latent Diffusion Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:51.541078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:51.541078Z digest=sha256:60a7f160e29338ba1d4bf214dccef3b006cd287cdbeaad19b524ecba97ed3e98

Observation 1a33835f-19c3-4a6d-b1dc-6dfd2f29a28e · outbound

This paper cites Label-Efficient Semantic Segmentation with Diffusion Models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Label-Efficient Semantic Segmentation with Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:51.657841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:51.657841Z digest=sha256:b72dfb74a8932379dceb8315d09d06e763508d5d74c251ac96b776f1f20a8aae

Observation 36b10831-eb2b-4b82-b2b1-97acb59f3dcb · outbound

This paper cites Unsupervised semantic correspondence using stable diffusion.Advances in Neural Information Processing Systems, 36:8266–8279, 2023.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Unsupervised semantic correspondence using stable diffusion.Advances in Neural Information Processing Systems, 36:8266–8279, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:57.299473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:51.784491Z digest=sha256:7cee6b2b14be998641a136c5991c69b6c5e4cb5792b1f7f0dec4914e99514ae2

Observation 89e1f53c-5f7d-4ac5-8f55-92c05ff80f58 · outbound

This paper cites Text-to-image diffusion models are zero shot classifiers.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Text-to-image diffusion models are zero shot classifiers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:57.108219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:51.910081Z digest=sha256:e3fb8c89e8e0f44fb793977d1311b2aefb61cd10403ebb6822a816236c649a22

Observation 9c4e2dac-5861-42d8-997c-587a4266f995 · outbound

This paper cites Diffusion model as representation learner.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Diffusion model as representation learner

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:52.029528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:52.029528Z digest=sha256:d0c2c0d42b3cd5227a6de3cf2702c3acc5183dd838fbe004cf357fdbc5a3c09c

Observation a90af4f5-6b70-4224-99f0-e9025975a2df · outbound

This paper cites Denoising diffusion autoencoders are unified self-supervised learners.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Denoising diffusion autoencoders are unified self-supervised learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:56.914753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:52.125636Z digest=sha256:6905238dd3f8cc230e29c78f19f872db9a07956fcc2f37064c953b0cdbc74be4

Observation ef68fce8-dcc5-4be5-ba3c-c8e514ffd416 · outbound

This paper cites Emer- gent correspondence from image diffusion.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Emer- gent correspondence from image diffusion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:56.716377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:52.237269Z digest=sha256:3f17d309934a9bf5c7949be1df76684432bf4495d6ef1cc6bcb46fc27a289f20

Observation d0699aa1-f24e-4543-b26f-b4d04a6e09c3 · outbound

This paper cites Clean- DIFT: Diffusion Features without Noise.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Clean- DIFT: Diffusion Features without Noise

Reference 29

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:35:54.529052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:52.371377Z digest=sha256:72755cf042b60f1924edcf18eaa45b84d2aa7681a914311d8c28b7bff30f2361

Observation ef7cd397-221e-46ab-bc58-3da91cf42e11 · outbound

This paper cites EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:52.518144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:52.518144Z digest=sha256:da341aac9ecdc978896a89e983a5b707faabf31332a21423109060f262afe711

Observation 8f161918-5091-4ebb-b847-f32c59fa174c · outbound

This paper cites What the DAAM: Interpreting Stable Diffusion Using Cross Attention.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback What the DAAM: Interpreting Stable Diffusion Using Cross Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:52.667595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:52.667595Z digest=sha256:9ac7d2c7ec966270ae91bc4c1db657a5b4e3e45cb13bc0805aa401a43c771cfe

Observation 470a2e14-431c-45b1-91fa-66608f9ded23 · outbound

This paper cites Your diffusion model is secretly a zero-shot classifier.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Your diffusion model is secretly a zero-shot classifier

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:52.769170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:52.769170Z digest=sha256:da881d341f3bcac7b702beeae0c25ce0660434e9e14596594c588a738b2531d0

Observation 8d1b51de-11a9-42b1-afbd-6b8f53c9c2bb · outbound

This paper cites Diffusiondet: Diffusion model for object detection.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Diffusiondet: Diffusion model for object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:56.522188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:52.872598Z digest=sha256:54bd13e49ec6073708b7008686cc974cb0d6b70c019fc871dd3462c76bcf3a0c

Observation 660801c3-b60b-4d54-871a-aeed6e9857b7 · outbound

This paper cites Readout guidance: Learning control from diffusion features.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Readout guidance: Learning control from diffusion features

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:56.319310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:52.975086Z digest=sha256:8a8568b8fb661636a58772ba7af65ead226efe194c9de6048ca425c07cfe2cf5

Observation bb614008-3af6-426d-87ec-9152aa2875e9 · outbound

This paper cites Diffusion hyperfeatures: Searching through time and space for semantic correspondence.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Diffusion hyperfeatures: Searching through time and space for semantic correspondence

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:56.135626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:53.098606Z digest=sha256:8dcf280d793faac60218866294c2b9a7cc614ff5fe523e8edc04588f7bfec435

Observation c2240105-1dcb-4fbb-be11-0800b15491af · outbound

This paper cites Distillation of diffusion features for semantic correspondence.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Distillation of diffusion features for semantic correspondence

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:55.952768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:53.228751Z digest=sha256:5bd0c9f59889bcb571e44900f5fc7cf7950da7c1c4803502bf51ca532d3cd95e

Observation ab039658-8d40-4ea9-be9b-00236b98d801 · outbound

This paper cites CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:53.348887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:53.348887Z digest=sha256:4e031b9ff3fbc8e83832069fbf696c96e38268689970d91a84d22eaa3aed38a2

Observation b4bb6863-4f77-4ba0-aa00-7063037e838a · outbound

This paper cites X-adapter: Adding universal compatibility of plugins for upgraded diffusion model.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback X-adapter: Adding universal compatibility of plugins for upgraded diffusion model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:55.803955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:53.442872Z digest=sha256:e947d7b84d8de715c2260d9cb8f335819bec98257108d8249381dcf0cfbc8c03

Observation 22be1095-d574-4829-91ac-756597ef4fbe · outbound

This paper cites Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:53.542608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:53.542608Z digest=sha256:07cb713f2afbf43fb4c135861a628752ad290e8f28014a1bbc4b3594f4f79904

Observation 7fec33fc-7041-455c-8ef4-505b94ee08a0 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Gligen: Open-set grounded text-to-image generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:55.660542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:53.639869Z digest=sha256:67cd2134779e8c244e66a692edc10c77bd5520c936e1e0bac59137217a7be886

Observation e7a4a8f5-3196-4144-a030-2a1b9c8f69b7 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:53.779487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:53.779487Z digest=sha256:90dba856163b98e429e60e75c1d1b54d13865ff4fbdec63cc7a0d7bdb8be06bc

Observation e5abe342-03dc-469b-93c5-04ee955b03dd · outbound

This paper cites Vision transformers for dense prediction.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Vision transformers for dense prediction

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:55.467397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:53.872465Z digest=sha256:bdeca57d6793448c098cbe550c4388783428d140bbe99c1f0517acafd9e52429

Observation 7634f2c2-fc5c-46a9-bbd7-2e61477a013f · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:54.012648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:35:54.012648Z digest=sha256:b708955047fb01f7b841ecaba195872130617fcdc5670191d9167d9b6f346cb2

Observation 0b0ae502-40eb-4ee8-8cc0-368f9b7345b0 · outbound

This paper cites Diffusers: State-of-the-art diffusion models.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Diffusers: State-of-the-art diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:55.293274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:54.092802Z digest=sha256:15140909fcad01d23df299eb8b643b76a1800ae20d17736b898dcc5b4d686f13

Observation 38e96513-a6ba-4463-b95f-78eaba3dd035 · outbound

This paper cites Deep residual learning for image recognition.

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback Deep residual learning for image recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:35:55.060581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T20:35:54.220222Z digest=sha256:f3e111c29ee0bb4b7b7ad358d6ba4bef5d1547b2cb77ab6f2d911430fd26774d

Pith citing papers

Observation b68a331c-fa61-4122-9cf5-aa4253609438 · inbound

Reflection Separation from a Single Image via Joint Latent Diffusion cites this paper.

Reflection Separation from a Single Image via Joint Latent Diffusion Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.322218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:47:22.111357Z digest=sha256:dc56a09d82638369eee9fd209f51eda6607c31b6bac73d9aae33ecd2ad7921c8