Pith. sign in

REVIEW 6 cited by

JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06347 v1 pith:CTD26GZV submitted 2023-10-10 cs.CV

JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling

classification cs.CV
keywords densediffusionjointnetbranchdistributiongenerationmodalitynetwork
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch. The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model. We demonstrate the effectiveness of JointNet by using RGBD diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGBD generation, dense depth prediction, depth-conditioned image generation, and coherent tile-based 3D panorama generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

    cs.CV 2026-07 conditional novelty 6.0

    A two-stage framework uses parallel depth generation and a 1M-sample dataset to improve geometric consistency in in-context panoramic image generation.

  2. Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

    cs.CV 2026-07 conditional novelty 6.0

    Geometry-aware RGB–depth pretraining plus a 1M multi-task panoramic dataset yields a unified in-context 360° generator with stronger FAED and seam consistency than prior methods.

  3. Modality Forcing for Scalable Spatial Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    Modality Forcing lets a single DiT produce image and depth outputs in any order after training on sparse real-world depth, with larger image-pretrained models yielding better depth accuracy and a 57% AbsRel reduction ...

  4. Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes

    cs.CV 2026-05 unverdicted novelty 5.0

    Introduces dual pose-image representation, cross-modal alignment, and iterative construction to improve prompt alignment and diversity in multi-person text-to-image generation.

  5. Enhancing Event-based Object Detection with Monocular Normal Maps

    cs.CV 2025-08 unverdicted novelty 5.0

    NRE-Net adds monocular normal maps as geometric priors to RGB-event fusion, delivering 3% AP50 gains on DSEC-Det-sub and PKU-DAVIS-SOD over dual-modal baselines.

  6. Enhancing Event-based Object Detection with Monocular Normal Maps

    cs.CV 2025-08 unverdicted novelty 5.0

    NRE-Net adds geometric priors from RGB-derived normal maps to RGB and event data via ADFM and EAFM fusion modules, reporting 3% AP50 gains over dual-modal baselines on driving datasets.