Pith. sign in

REVIEW 2 major objections 2 minor 65 cited by

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

T0 review · 2 major / 2 minor · reviewed 2026-05-20 · grok-4.3

Pith's one-line read DragNUWA achieves fine-grained control in open-domain video generation by integrating text, image, and trajectory information.

desk verdict DragNUWA adds practical text-image-trajectory control to diffusion video models via three targeted modules, but the superiority claim needs concrete numbers to land. read the letter →

arxiv 2308.08089 v1 pith:UDEIT65R submitted 2023-08-16 cs.CV

classification cs.CV
keywords videogenerationdiffusionmodelstrajectorycontrolfine-grainedopen-domainmultimodalconditioningmotionguidance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces DragNUWA, a diffusion-based model for video generation that accepts text, image, and trajectory inputs together. Text supplies semantic guidance, the image fixes spatial layout, and the trajectory dictates motion paths. Earlier methods handled only one of these signals or worked only on simple datasets like Human3.6M, which restricted their use on real scenes and complex motions. DragNUWA adds a Trajectory Sampler to accept arbitrary curves on any image, Multiscale Fusion to blend control at different resolutions, and an Adaptive Training procedure to keep generated frames consistent with the inputs. A sympathetic reader would see this as a step toward video creation tools that let users direct content, placement, and movement with combined instructions.

What carries the argument

Trajectory Sampler for open-domain arbitrary trajectories, Multiscale Fusion for varying control granularities, and Adaptive Training for motion consistency, all combined with text and image conditioning in a diffusion video model.

What would settle it

Generate videos from complex curved trajectories drawn on diverse real-world open-domain images and measure whether the motion paths are followed accurately while content and semantics remain stable and free of visible artifacts.

Watch

Extended reading notes

Core claim

DragNUWA is an open-domain diffusion-based video generation model that simultaneously introduces text, image, and trajectory information to provide fine-grained control over video content from semantic, spatial, and temporal perspectives. It resolves limited open-domain trajectory control by proposing a Trajectory Sampler to enable arbitrary trajectories, Multiscale Fusion to control trajectories at different granularities, and an Adaptive Training strategy to generate consistent videos that follow the trajectories.

Load-bearing premise

The Trajectory Sampler, Multiscale Fusion, and Adaptive Training can reliably produce consistent videos that follow arbitrary complex curved trajectories on open-domain images without motion artifacts or semantic drift.

Editorial extensions

If this is right

  • Videos can be created that follow user-specified arbitrary curved trajectories overlaid on any input image.
  • Semantic content from text, spatial details from the image, and temporal motion from the trajectory are controlled at the same time.
  • The model handles open-domain scenes rather than being limited to narrow datasets like Human3.6M.
  • Generated sequences maintain consistency with the trajectory across frames.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same multi-signal conditioning pattern could be tested on related tasks such as image animation or 3D scene generation.
  • Interactive interfaces might let users sketch paths directly on an image to define desired motion.
  • The method could reduce reliance on single-modality training data when scaling controllable generation systems.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper introduces DragNUWA, an open-domain diffusion-based video generation model that integrates text, image, and trajectory inputs to enable fine-grained control over video content from semantic, spatial, and temporal perspectives. It proposes a Trajectory Sampler (TS) to support arbitrary trajectories on open-domain images, Multiscale Fusion (MF) to handle trajectories at varying granularities, and an Adaptive Training (AT) strategy to produce consistent videos that follow the specified trajectories. The central claim is that this combination yields superior performance in fine-grained controllable video generation compared to prior single-modality or limited-domain approaches, with experiments asserted to validate the effectiveness.

Significance. If the results hold, the work would advance multi-modal controllable video generation by addressing the limitations of single-modality control and restriction to simple datasets such as Human3.6M. Enabling open-domain handling of complex curved trajectories via the proposed TS, MF, and AT components could support more precise applications in animation and content creation. The structured decomposition of trajectory modeling provides a clear technical contribution to the diffusion conditioning literature.

major comments (2)
  1. [Experiments] Experiments section: the abstract asserts experimental validation and superior performance, yet no quantitative metrics (e.g., FID, FVD, or user-study scores), dataset details, or ablation results on TS/MF/AT are provided in the summary; without these the central claim of effectiveness rests on an unverified assertion and requires explicit tables comparing against baselines on open-domain data.
  2. [Method] Method section, description of Adaptive Training (AT): the strategy is presented as ensuring consistency and avoiding motion artifacts or semantic drift for arbitrary curved trajectories, but the concrete loss formulation, sampling schedule, or conditioning weight schedule is not specified; this leaves the load-bearing claim that AT reliably produces artifact-free output ungrounded in the provided equations or pseudocode.
minor comments (2)
  1. The homepage link is given but the manuscript does not include a direct pointer to the released code or model weights, which would aid reproducibility.
  2. [Method] Notation for the three conditioning modalities (text, image, trajectory) should be introduced with explicit symbols in the method overview to improve clarity when describing the fusion step.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their detailed and constructive comments. We address each major comment point by point below and outline the revisions we will make to strengthen the manuscript.

read point-by-point responses
  1. Referee: [Experiments] Experiments section: the abstract asserts experimental validation and superior performance, yet no quantitative metrics (e.g., FID, FVD, or user-study scores), dataset details, or ablation results on TS/MF/AT are provided in the summary; without these the central claim of effectiveness rests on an unverified assertion and requires explicit tables comparing against baselines on open-domain data.

    Authors: We appreciate the referee highlighting the need for clearer quantitative support. The full manuscript contains experimental results including user studies demonstrating superior performance; however, to directly address this concern, we will add explicit tables in the revised version reporting quantitative metrics (such as FID and FVD where relevant), dataset details, and ablation studies isolating the contributions of TS, MF, and AT, with direct comparisons to baselines on open-domain data. revision: yes

  2. Referee: [Method] Method section, description of Adaptive Training (AT): the strategy is presented as ensuring consistency and avoiding motion artifacts or semantic drift for arbitrary curved trajectories, but the concrete loss formulation, sampling schedule, or conditioning weight schedule is not specified; this leaves the load-bearing claim that AT reliably produces artifact-free output ungrounded in the provided equations or pseudocode.

    Authors: We agree that the Adaptive Training description would benefit from greater specificity. In the revised manuscript, we will include the concrete loss formulation for AT, along with the sampling schedule and conditioning weight schedule. These additions will better substantiate the claims regarding consistency and the avoidance of motion artifacts or semantic drift for arbitrary trajectories. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity detected

full rationale

The paper proposes DragNUWA as a diffusion-based architecture that adds text, image, and trajectory conditioning, along with three explicitly defined new components (Trajectory Sampler for open-domain paths, Multiscale Fusion for granularity, and Adaptive Training for consistency). These elements are introduced to address stated limitations in prior work and are described directly in the method without any equations or claims that reduce the performance gains to a fitted parameter, self-definition, or self-citation chain. The derivation remains self-contained as a coherent extension of standard diffusion conditioning, with validation left to experiments rather than internal reduction.

Assumptions & free parameters 1 free parameters · 1 assumptions · 3 invented entities

The proposal rests on standard diffusion-model conditioning assumptions plus three newly introduced modules whose effectiveness is asserted rather than derived from first principles.

free parameters (1)
  • diffusion conditioning weights for text/image/trajectory
    Standard learned or hand-tuned scalars that balance the three control signals during sampling.
assumptions (1)
  • domain assumption Diffusion models can be jointly conditioned on semantic, spatial, and temporal signals without destructive interference.
    Invoked when the paper states that simultaneous introduction of the three inputs yields fine-grained control.
invented entities (3)
  • Trajectory Sampler (TS)
    purpose: Enable open-domain control of arbitrary trajectories
    New module introduced to sample points along user-drawn paths on complex scenes.
  • Multiscale Fusion (MF)
    purpose: Control trajectories at different granularities
    New fusion mechanism to combine trajectory information across scales.
  • Adaptive Training (AT) strategy
    purpose: Generate consistent videos following trajectories
    New training procedure claimed to enforce trajectory adherence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory." pith.science (2026). https://pith.science/paper/UDEIT65R

@misc{pith2026230808089,
  author       = {Pith},
  title        = {Pith review of: DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UDEIT65R}},
  note         = {Machine review of arXiv:2308.08089}
}
read the original abstract

Controllable video generation has gained significant attention in recent years. However, two main limitations persist: Firstly, most existing works focus on either text, image, or trajectory-based control, leading to an inability to achieve fine-grained control in videos. Secondly, trajectory control research is still in its early stages, with most experiments being conducted on simple datasets like Human3.6M. This constraint limits the models' capability to process open-domain images and effectively handle complex curved trajectories. In this paper, we propose DragNUWA, an open-domain diffusion-based video generation model. To tackle the issue of insufficient control granularity in existing works, we simultaneously introduce text, image, and trajectory information to provide fine-grained control over video content from semantic, spatial, and temporal perspectives. To resolve the problem of limited open-domain trajectory control in current research, We propose trajectory modeling with three aspects: a Trajectory Sampler (TS) to enable open-domain control of arbitrary trajectories, a Multiscale Fusion (MF) to control trajectories in different granularities, and an Adaptive Training (AT) strategy to generate consistent videos following trajectories. Our experiments validate the effectiveness of DragNUWA, demonstrating its superior performance in fine-grained control in video generation. The homepage link is \url{https://www.microsoft.com/en-us/research/project/dragnuwa/}

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 65 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 65 Pith citations

  1. QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    QWERTY enables training-free motion control in pretrained image-to-video DiTs by warping the frame-invariant semantic subspace of queries in 3D full attention and using the predicted noise as self-guidance for latent ...

  2. SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    SVI-Bench provides 35K hours of sports video with 9 tasks across four cognitive levels, revealing models drop from ~74% on action QA to 5% on agentic evidence integration.

  3. CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    CoMoGen generates controllable interactive video from mask sequences and images by encoding masks into MMDiT via MaskAdapter and LoRA on motion layers, claiming SOTA motion fidelity.

  4. MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    MotiMotion adds visual reasoning via a training-free VLM to refine primary trajectories and hallucinate secondary motions, plus a confidence-aware guidance scheme, yielding more plausible interactions on the new MotiB...

  5. Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    PREX decomposes target 4D video volumes into Preserve, Reveal, and Expand roles with a region-aware adapter on a frozen diffusion backbone, trained via proxy tasks, and introduces the PREBench benchmark to reduce regi...

  6. Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Introduces Eulerian motion guidance with bidirectional geometric consistency to improve training speed and temporal quality in diffusion-based image animation.

  7. ActionParty: Multi-Subject Action Binding in Generative Video Games

    cs.CV 2026-04 conditional novelty 7.0 of 10

    ActionParty binds discrete actions to individual subjects in a single generated video by jointly modeling subject state tokens and video latents, controlling up to seven players across 46 Melting Pot games.

  8. GPS as a Control Signal for Image Generation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.

  9. Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion-based image animation method that uses rendered colored spheres and a checkerboard world envelope as 3D-aware motion control signals, enabling joint camera and object motion control with fine granularity.

  10. FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single video-diffusion framework composites both static images and dynamic footage along user-defined trajectories by transporting canonical foreground latents directly into the background latent sequence.

  11. Robot-Factored World Models via Robot Rendering

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Conditioning a video world model on rendered nominal robot trajectories (URDF mesh + depth) instead of raw actions or logged future states improves action-following and enables zero-shot embodiment change.

  12. GraphVid: Interactive Graph-Controllable Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GraphVid controls video generation with user-editable interaction scene graphs, reporting FID/FVD improvements over trajectory- and text-physics baselines using 0.6B trainable parameters.

  13. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    MultiRef-Compass is a 350-sample benchmark and 14-metric protocol for multi-reference-to-audio-video generation; current models still fail at reference binding and audio-visual consistency.

  14. Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    UniCaMo builds 3D-grounded motion-consistent input noise from sparse 3D tracks and sphere-sampled noise so pretrained video diffusion models jointly control object and camera motion without architectural changes.

  15. HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    HandsOnWorld creates a hand-controlled egocentric video generator from unconstrained monocular video via a new EgoVid-Pro dataset from monocular reconstruction and a Plücker Hand Map that disentangles camera and hand motion.

  16. In-context Region-based Drag: Drag Any Region to Any Shape

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ICRDrag performs region-based drag editing in diffusion models by feeding source image, source mask, and target mask into an in-context framework with image-mask attention consistency and source-target attention corre...

  17. OrthoMotion:Disentangling Camera and Subject Motion via Geometry Semantics Orthogonal Attention

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    OrthoMotion disentangles camera and subject motion in video generation by splitting attention into algebraically complementary geometric (RoPE rotation) and semantic (gated value) channels driven to orthogonality by a...

  18. OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    OmniDirector introduces a grid-based camera representation and hierarchical prompt agent for multi-shot camera cloning in video diffusion models trained on million-scale unpaired data.

  19. Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    T2Mo generates controllable dynamic 3D shapes by conditioning on both text semantics and 3D trajectories with a shape-grounded embedding for arbitrary inputs.

  20. Compositional Video Generation via Inference-Time Guidance

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CVG improves compositional faithfulness in frozen text-to-video diffusion models by steering early denoising steps with gradients from a classifier trained on the model's own cross-attention features.

  21. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.

  22. GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geom...

  23. TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    TransFlow generates 218,008 training triplets by animating DUTS images with Stable Video Diffusion and estimating optical flow with RAFT, then uses them to train a two-stream video SOD network that improves S-measure ...

  24. MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MotionShot transfers motion from a reference video to an unseen target object in text-to-video generation by combining semantic and morphological alignment in a training-free pipeline.

  25. AnyI2V: Animating Any Conditional Image with Motion Control

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.

  26. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

  27. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.

  28. IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model

    cs.GR 2025-05 conditional novelty 6.0 of 10

    A motion diffusion model with decoupled trajectory and keyframe-pose control, wrapped in an MLLM agent system, produces more controllable 3D human motion from images and text.

  29. MotionPro: A Precise Motion Controller for Image-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MotionPro uses region-wise trajectories and a motion mask to control object and camera motion in image-to-video generation, reporting improved trajectory alignment over prior methods.

  30. Hybrid Neural-MPM for Interactive Fluid Simulations in Real-Time

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A hybrid neural-MPM solver with a chaos-triggered fallback and a diffusion-based sketch controller enables real-time interactive fluid simulation with user control.

  31. LMP: Leveraging Motion Prior in Zero-Shot Video Generation with Diffusion Transformer

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LMP transfers motion from a reference video to newly generated videos in text-to-video and image-to-video settings without training, using attention maps in a frozen diffusion transformer.

  32. FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A two-stage video diffusion pipeline uses optical flow as its control signal, achieving accurate camera control and natural object motion without ground-truth camera-parameter labels during training.

  33. VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    VidCRAFT3 is a single image-to-video diffusion system that accepts camera, object, and lighting direction controls separately or jointly, trained in three stages with a new synthetic lighting dataset.

  34. OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    OmniPhysGS lets each Gaussian in a 3D scene select from 12 expert material models, supervised by a text-to-video diffusion model, to generate dynamics for multiple materials.

  35. PreciseCam: Precise Camera Control for Text-to-Image Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PreciseCam enables precise camera control (roll, pitch, vFoV, distortion) in text-to-image generation by conditioning SDXL with Perspective Field maps and a new dataset of 57,380 images.

  36. Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Ouroboros-Diffusion improves long video consistency by combining low-frequency tail noise, subject-aware cross-frame attention, and self-recurrent gradient guidance in a tuning-free FIFO diffusion queue.

  37. Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A training-free method that matches sparse-point inter-frame feature correlations to transfer reference motion to generated videos with improved temporal consistency.

  38. TransPixeler: Advancing Text-to-Video Generation with Transparency

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A LoRA-based adaptation of DiT video generators that jointly outputs aligned RGB and alpha channels via extra tokens, shared positions, and attention masking.

  39. LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LeviTor controls 3D object trajectories in generated videos by feeding K-means clustered mask points with estimated depth into a video diffusion model.

  40. Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A fine-tuned Llama 3 LLM writes discrete human pose tokens from a text prompt, and a pose-conditioned diffusion video renderer turns them into videos, improving human motion fidelity.

  41. MotionBridge: Dynamic Video Inbetweening with Flexible Controls

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MotionBridge generates interpolated video frames between two images while following user-supplied trajectory, mask, keyframe, guide-pixel, and text controls.

  42. InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    InterDyn fine-tunes Stable Video Diffusion with a ControlNet-style branch so that a hand-mask control signal drives plausible, temporally consistent videos of object interactions.

  43. OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A drag-style motion control method for 360 degree image-to-video generation, built on spherical trajectory estimation and joint fine-tuning of a pretrained video diffusion model.

  44. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.

  45. ObjCtrl-2.5D: Training-free Object Control with Camera Poses

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free method that lifts 2D object trajectories into camera poses with depth and uses a frozen camera-control video model to achieve more accurate and 3D-aware object motion, including rotation.

  46. HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HANDI generates hand-centric videos from an image and text prompt via automatic motion-area localization and a hand refinement loss.

  47. InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A three-stage pipeline generates up to 100,000 square meters of dynamic 3D driving scenes with 200-frame videos, controlled by HD maps, bounding boxes, and text.

  48. Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Sampling-time energy guidance steers a frozen rectified-flow driving world model's ego trajectory to a braking target, but the generated video does not follow under current joint self-attention.

  49. WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    A video world model framework that uses LLM-orchestrated 3D trajectories as control signals for generation to achieve persistent dynamic object memory and viewpoint freedom.

  50. Perceptual 3D Simulation With Physical World Modeling

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    P3Sim integrates a probabilistic physical world model with geometric conditioning and persistent memory to simulate 3D scenes under partial observations and incomplete transforms.

  51. WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    WorldCraft introduces NWT, SP-LoRA, and TASP to enable object trajectory control in video-based world models while preserving camera navigation.

  52. Unified 3D Scene Understanding Through Physical World Modeling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A probabilistic graphical model called 3WM unifies 3D vision tasks into one system that performs them zero-shot by selecting different inference pathways through multimodal scene nodes.

  53. DriveCtrl: Conditioned Sim-to-Real Driving Video Generation

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    DriveCtrl is a depth-conditioned controllable framework that generates realistic driving videos from simulation while preserving annotations and scene dynamics.

  54. Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Tora2 adds decoupled personalization embeddings, gated self-attention binding, and contrastive learning to Tora, enabling simultaneous appearance and trajectory customization for multiple entities in generated video.

  55. LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LiON-LoRA adds a learned scaling token to video-diffusion LoRA adapters, enabling linear and independent control of camera trajectory and object motion strength.

  56. ATI: Any Trajectory Instruction for Controllable Video Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ATI injects user-drawn point trajectories as soft Gaussian feature masks into a pretrained image-to-video diffusion model, enabling unified camera, object, and local motion control.

  57. RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Metric-scale depth alignment plus scene-constrained noise shaping improves camera controllability and video quality for image-to-video generation on RealEstate10K.

  58. PhysAnimator: Physics-Guided Generative Cartoon Animation

    cs.GR 2025-01 conditional novelty 5.0 of 10

    PhysAnimator combines 2D deformable-body physics simulation with a sketch-guided video diffusion model to animate static anime illustrations with controllable, physically plausible motion.

  59. X-Dyna: Expressive Dynamic Human Image Animation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A diffusion-based pipeline that animates a single human image with pose, expression, and dynamic background effects from a driving video, outperforming prior methods on dynamic detail metrics.

  60. UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Training small adapters on static videos, then applying them at low intensity, improves the consistency and frame quality of diffusion-based video generators without retraining the base model.

See all 65 Pith citations

Reference graph

Works this paper leans on

297 extracted references · 297 canonical work pages · cited by 65 Pith papers (see all)

  1. [1]

    Click To Move : Controlling Video Generation With Sparse Motion

    Pierfrancesco Ardino, Marco De Nadai, Bruno Lepri, Elisa Ricci, and St \'e phane Lathuili \`e re. Click To Move : Controlling Video Generation With Sparse Motion . In Proceedings of the IEEE / CVF International Conference on Computer Vision , pp.\ 14749--14758, 2021

  2. [2]

    Frozen in time: A joint video and image encoder for end-to-end retrieval

    Max Bain, Arsha Nagrani, G \"u l Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE / CVF International Conference on Computer Vision , pp.\ 1728--1738, 2021

  3. [3]

    Ipoke: Poking a still image for controlled stochastic video synthesis

    Andreas Blattmann, Timo Milbich, Michael Dorkenwald, and Bj \"o rn Ommer. Ipoke: Poking a still image for controlled stochastic video synthesis. In Proceedings of the IEEE / CVF International Conference on Computer Vision , pp.\ 14707--14717, 2021 a

  4. [4]

    Understanding Object Dynamics for Interactive Image-to-Video Synthesis

    Andreas Blattmann, Timo Milbich, Michael Dorkenwald, and Bjorn Ommer. Understanding Object Dynamics for Interactive Image-to-Video Synthesis . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pp.\ 5171--5181, 2021 b

  5. [5]

    Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A. Efros. Everybody Dance Now . In Proceedings of the IEEE / CVF International Conference on Computer Vision , pp.\ 5933--5942, 2019

  6. [7]

    Recurrent Environment Simulators

    Silvia Chiappa, S \'e bastien Racaniere, Daan Wierstra, and Shakir Mohamed. Recurrent Environment Simulators . In International Conference on Learning Representations , November 2016

  7. [9]

    Controllable Video Generation With Sparse Trajectories

    Zekun Hao, Xun Huang, and Serge Belongie. Controllable Video Generation With Sparse Trajectories . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pp.\ 7854--7863, 2018

  8. [10]

    Imagen Video: High Definition Video Generation with Diffusion Models

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, and David J. Fleet. Imagen video: High video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022

Show all 297 references
  1. [11]

    CogVideo : Large-scale Pretraining for Text-to-Video Generation via Transformers

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. CogVideo : Large-scale Pretraining for Text-to-Video Generation via Transformers . arXiv preprint arXiv:2205.15868, 2022

  2. [12]

    Make It Move : Controllable Image-to-Video Generation With Text Descriptions

    Yaosi Hu, Chong Luo, and Zhenzhong Chen. Make It Move : Controllable Image-to-Video Generation With Text Descriptions . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pp.\ 18219--18228, 2022

  3. [13]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014

  4. [14]

    BLIP-2 : Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2 : Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models . arXiv preprint arXiv:2301.12597, 2023

  5. [15]

    Xiaodan Liang, Lisa Lee, Wei Dai, and Eric P. Xing. Dual Motion GAN for Future-Flow Embedded Video Prediction . In Proceedings of the IEEE International Conference on Computer Vision , pp.\ 1744--1752, 2017

  6. [16]

    Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning

    William Lotter, Gabriel Kreiman, and David Cox. Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning . In International Conference on Learning Representations , November 2016

  7. [17]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, and Jack Clark. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning , pp.\...

  8. [18]

    High- Resolution Image Synthesis With Latent Diffusion Models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High- Resolution Image Synthesis With Latent Diffusion Models . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pp.\ 10684--10695, 2022

  9. [19]

    Make- A-Video : Text-to-Video Generation without Text-Video Data , September 2022

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make- A-Video : Text-to-Video Generation without Text-Video Data , September 2022

  10. [20]

    Unsupervised Learning of Video Representations using LSTMs

    Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov. Unsupervised Learning of Video Representations using LSTMs . In Proceedings of the 32nd International Conference on Machine Learning , pp.\ 843--852. PMLR , June 2015

  11. [21]

    Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions

    Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions . In ICLR , September 2022

  12. [22]

    The Pose Knows : Video Forecasting by Generating Pose Futures

    Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert. The Pose Knows : Video Forecasting by Generating Pose Futures . In Proceedings of the IEEE International Conference on Computer Vision , pp.\ 3332--3341, 2017

  13. [23]

    Deep High-Resolution Representation Learning for Visual Recognition

    Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu, and Bin Xiao. Deep High-Resolution Representation Learning for Visual Recognition . IEEE Transactions on Pattern Analysis and Machine Intell...

  14. [24]

    Few-shot video-to-video synthesis

    Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Jan Kautz, and Bryan Catanzaro. Few-shot video-to-video synthesis. In Proceedings of the 33rd International Conference on Neural Information Processing Systems , pp.\ 5013--5024, Red Hook, NY, USA , December 2019. Curran Ass...

  15. [25]

    VideoComposer : Compositional Video Synthesis with Motion Controllability , June 2023

    Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. VideoComposer : Compositional Video Synthesis with Motion Controllability , June 2023

  16. [26]

    Hierarchical Long-term Video Prediction without Supervision

    Nevan Wichers, Ruben Villegas, Dumitru Erhan, and Honglak Lee. Hierarchical Long-term Video Prediction without Supervision . In Proceedings of the 35th International Conference on Machine Learning , pp.\ 6038--6046. PMLR , July 2018

  17. [27]

    GODIVA : Generating Open-DomaIn Videos from nAtural Descriptions

    Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. GODIVA : Generating Open-DomaIn Videos from nAtural Descriptions . arXiv:2104.14806 [cs], April 2021

  18. [28]

    N " UWA : Visual Synthesis Pre-training for Neural visUal World creAtion

    Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. N " UWA : Visual Synthesis Pre-training for Neural visUal World creAtion . In Proceedings of the European Conference on Computer Vision ( ECCV ) , 2022

  19. [29]

    Future Video Synthesis With Object Motion Prediction

    Yue Wu, Rongrong Gao, Jaesik Park, and Qifeng Chen. Future Video Synthesis With Object Motion Prediction . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pp.\ 5539--5548, 2020

  20. [30]

    Unifying flow, stereo and depth estimation

    Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying flow, stereo and depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  21. [31]

    NUWA-XL : Diffusion over Diffusion for eXtremely Long Video Generation

    Shengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, and Fan Yang. NUWA-XL : Diffusion over Diffusion for eXtremely Long Video Generation . arXiv preprint arXiv:2303.12346, 2023

  22. [32]

    DTVNet : Dynamic Time-Lapse Video Generation via Single Still Image

    Jiangning Zhang, Chao Xu, Liang Liu, Mengmeng Wang, Xia Wu, Yong Liu, and Yunliang Jiang. DTVNet : Dynamic Time-Lapse Video Generation via Single Still Image . In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (eds.), Computer Vision ECCV 2020 , Lecture Note...

  23. [33]

    IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

    Unifying Flow, Stereo and Depth Estimation , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

  24. [34]

    99firms.com , urldate =

    2019's. 99firms.com , urldate =

  25. [35]

    2019 , month = may, journal =

    Around 40,000 Songs Are Uploaded to. 2019 , month = may, journal =

  26. [36]

    2018 , pages =

    Watch. 2018 , pages =

  27. [37]

    1995 , pages =

    Financial Applications of Learning from Hints , booktitle =. 1995 , pages =

  28. [38]

    1995 , journal =

    Hints , author =. 1995 , journal =

  29. [39]

    1993 , pages =

    A Method for Learning from Hints , booktitle =. 1993 , pages =

  30. [40]

    2018 , journal =

    Towards High Resolution Video Generation with Progressive Growing of Sliced Wasserstein Gans , author =. 2018 , journal =. 1810.02419 , archiveprefix =

  31. [41]

    arXiv:1810.12440 [cs] , eprint =

    Acharya, Manoj and Kafle, Kushal and Kanan, Christopher , year =. arXiv:1810.12440 [cs] , eprint =

  32. [42]

    Explicit

    Aditya, Somak and Yang, Yezhou and Baral, Chitta , year =. Explicit

  33. [43]

    Question

    Adiwardana, Daniel De Freitas and Shakeri, Siamak , urldate =. Question

  34. [44]

    Analyzing the

    Agrawal, Aishwarya and Batra, Dhruv and Parikh, Devi , year =. Analyzing the

  35. [45]

    Agrawal, Aishwarya and Kembhavi, Aniruddha and Batra, Dhruv and Parikh, Devi , year =. C-. arXiv:1704.08243 [cs] , eprint =

  36. [46]

    Agrawal, Aishwarya and Batra, Dhruv and Parikh, Devi and Kembhavi, Aniruddha , year =. Don't. arxiv , file =:1712.00377 , urldate =

  37. [47]

    Lawrence and Parikh, Devi and Batra, Dhruv , year =

    Agrawal, Aishwarya and Lu, Jiasen and Antol, Stanislaw and Mitchell, Margaret and Zitnick, C. Lawrence and Parikh, Devi and Batra, Dhruv , year =. International Journal of Computer Vision , volume =. doi:10.1007/s11263-016-0966-6 , urldate =

  38. [48]

    Scale-Space Flow for End-to-End Optimized Video Compression , booktitle =

    Agustsson, Eirikur and Minnen, David and Johnston, Nick and Balle, Johannes and Hwang, Sung Jin and Toderici, George , year =. Scale-Space Flow for End-to-End Optimized Video Compression , booktitle =

  39. [49]

    2016 , journal =

    A Neural Knowledge Language Model , author =. 2016 , journal =. arxiv , file =:1608.00318 , urldate =

  40. [50]

    arXiv preprint arXiv:2004.08483 , eprint =

    Ainslie, Joshua and Ontanon, Santiago and Alberti, Chris and Pham, Philip and Ravula, Anirudh and Sanghai, Sumit , year =. arXiv preprint arXiv:2004.08483 , eprint =

  41. [51]

    Proceedings of the

    Akan, Adil Kaan and Erdem, Erkut and Erdem, Aykut and G. Proceedings of the. 2021 , pages =

  42. [52]

    2022 , journal =

    Stochastic Video Prediction with Structure and Motion , author =. 2022 , journal =. 2203.10528 , archiveprefix =

  43. [53]

    Contextual

    Akbik, Alan and Blythe, Duncan and Vollgraf, Roland , year =. Contextual

  44. [54]

    Flamingo: A

    Alayrac, Jean-Baptiste and Donahue, Jeff and Luc, Pauline and Miech, Antoine and Barr, Iain and Hasson, Yana and Lenc, Karel and Mensch, Arthur and Millican, Katie and Reynolds, Malcolm and Ring, Roman and Rutherford, Eliza and Cabi, Serkan and Han, Tengda and Gong, Zhitao and...

  45. [55]

    2017 , journal =

    Learning from Narrated Instruction Videos , author =. 2017 , journal =

  46. [56]

    Alayrac, Jean-Baptiste and Recasens, Adri. Self-. 2020 , journal =. 2006.16228 , archiveprefix =

  47. [57]

    Albanie, Samuel and Liu, Yang and Nagrani, Arsha and Miech, Antoine and Coto, Ernesto and Laptev, Ivan and Sukthankar, Rahul and Ghanem, Bernard and Zisserman, Andrew and Gabeur, Valentin , year =. The. arXiv preprint arXiv:2008.00744 , eprint =

  48. [58]

    2019 , journal =

    Fusion of Detected Objects in Text for Visual Question Answering , author =. 2019 , journal =. 1908.05054 , archiveprefix =

  49. [59]

    Applications of Generative Adversarial Networks (Gans):

    Alqahtani, Hamed and. Applications of Generative Adversarial Networks (Gans):. 2021 , journal =

  50. [60]

    Bottom-up and Top-down Attention for Image Captioning and Visual Question Answering , booktitle =

    Anderson, Peter and He, Xiaodong and Buehler, Chris and Teney, Damien and Johnson, Mark and Gould, Stephen and Zhang, Lei , year =. Bottom-up and Top-down Attention for Image Captioning and Visual Question Answering , booktitle =

  51. [61]

    Anderson, Peter and Fernando, Basura and Johnson, Mark and Gould, Stephen , year =. Spice:. European

  52. [62]

    Learning to Compose Neural Networks for Question Answering , booktitle =

    Andreas, Jacob and Rohrbach, Marcus and Darrell, Trevor and Klein, Dan , year =. Learning to Compose Neural Networks for Question Answering , booktitle =

  53. [63]

    Neural Module Networks , booktitle =

    Andreas, Jacob and Rohrbach, Marcus and Darrell, Trevor and Klein, Dan , year =. Neural Module Networks , booktitle =

  54. [64]

    Relationships from

    Andrews, Martin and AI, Red Dragon and Witteveen, Sam , keywords =. Relationships from

  55. [65]

    and Parikh, Devi , year =

    Antol, Stanislaw and Agrawal, Aishwarya and Lu, Jiasen and Mitchell, Margaret and Batra, Dhruv and Lawrence Zitnick, C. and Parikh, Devi , year =. Vqa:

  56. [66]

    Ardino, Pierfrancesco and De Nadai, Marco and Lepri, Bruno and Ricci, Elisa and Lathuili. Click. Proceedings of the. 2021 , pages =

  57. [67]

    Arnab, Anurag and Dehghani, Mostafa and Heigold, Georg and Sun, Chen and Lu. Vivit:. Proceedings of the. 2021 , pages =

  58. [68]

    Variational Transformer Networks for Layout Generation , booktitle =

    Arroyo, Diego Martin and Postels, Janis and Tombari, Federico , year =. Variational Transformer Networks for Layout Generation , booktitle =

  59. [69]

    Avrahami, Omri and Lischinski, Dani and Fried, Ohad , year =. Blended. arXiv preprint arXiv:2111.14818 , eprint =

  60. [70]

    Avrahami, Omri and Fried, Ohad and Lischinski, Dani , year =. Blended. doi:10.48550/arXiv.2206.02779 , urldate =. arxiv , file =:2206.02779 , primaryclass =

  61. [71]

    2017 , journal =

    Stochastic Variational Video Prediction , author =. 2017 , journal =. 1710.11252 , archiveprefix =

  62. [72]

    and Levine, Sergey , year =

    Babaeizadeh, Mohammad and Finn, Chelsea and Erhan, Dumitru and Campbell, Roy H. and Levine, Sergey , year =. Stochastic

  63. [73]

    2014 , journal =

    Neural Machine Translation by Jointly Learning to Align and Translate , author =. 2014 , journal =. arxiv , file =:1409.0473 , urldate =

  64. [74]

    doi:10.48550/arXiv.2206.14797 , urldate =

    Bahmani, Sherwin and Park, Jeong Joon and Paschalidou, Despoina and Tang, Hao and Wetzstein, Gordon and Guibas, Leonidas and Van Gool, Luc and Timofte, Radu , year =. doi:10.48550/arXiv.2206.14797 , urldate =. arxiv , keywords =:2206.14797 , primaryclass =

  65. [75]

    航空计算技术 , volume =

    白, 林亭 and 文, 鹏程 and 李, 亚晖 , year =. 航空计算技术 , volume =

  66. [76]

    Frozen in Time:

    Bain, Max and Nagrani, Arsha and Varol, G. Frozen in Time:. Proceedings of the. 2021 , pages =

  67. [77]

    Conditional

    Balaji, Yogesh and Min, Martin Renqiang and Bai, Bing and Chellappa, Rama and Graf, Hans Peter , year =. Conditional

  68. [78]

    arXiv preprint arXiv:2211.01324 , eprint =

    Balaji, Yogesh and Nah, Seungjun and Huang, Xun and Vahdat, Arash and Song, Jiaming and Kreis, Karsten and Aittala, Miika and Aila, Timo and Laine, Samuli and Catanzaro, Bryan , year =. arXiv preprint arXiv:2211.01324 , eprint =

  69. [79]

    and Kazemi, Hamid and Huang, Furong and Goldblum, Micah and Geiping, Jonas and Goldstein, Tom , year =

    Bansal, Arpit and Borgnia, Eitan and Chu, Hong-Min and Li, Jie S. and Kazemi, Hamid and Huang, Furong and Goldblum, Micah and Geiping, Jonas and Goldstein, Tom , year =. Cold Diffusion:. arXiv preprint arXiv:2208.09392 , eprint =

  70. [80]

    Analytic-

    Bao, Fan and Li, Chongxuan and Zhu, Jun and Zhang, Bo , year =. Analytic-. International

  71. [81]

    Bao, Fan and Nie, Shen and Xue, Kaiwen and Li, Chongxuan and Pu, Shi and Wang, Yaole and Yue, Gang and Cao, Yue and Su, Hang and Zhu, Jun , year =. One. arXiv preprint arXiv:2303.06555 , eprint =

  72. [82]

    Bao, Fan and Nie, Shen and Xue, Kaiwen and Li, Chongxuan and Pu, Shi and Wang, Yaole and Yue, Gang and Cao, Yue and Su, Hang and Zhu, Jun , year =. One. doi:10.48550/arXiv.2303.06555 , urldate =. arxiv , file =:2303.06555 , primaryclass =

  73. [83]

    Bao, Hangbo and Dong, Li and Wei, Furu and Wang, Wenhui and Yang, Nan and Liu, Xiaodong and Wang, Yu and Piao, Songhao and Gao, Jianfeng and Zhou, Ming and Hon, Hsiao-Wuen , year =

  74. [84]

    and Mildenhall, Ben and Verbin, Dor and Srinivasan, Pratul P

    Barron, Jonathan T. and Mildenhall, Ben and Verbin, Dor and Srinivasan, Pratul P. and Hedman, Peter , year =. Mip-Nerf 360:. Proceedings of the

  75. [85]

    and Mildenhall, Ben and Verbin, Dor and Srinivasan, Pratul P

    Barron, Jonathan T. and Mildenhall, Ben and Verbin, Dor and Srinivasan, Pratul P. and Hedman, Peter , year =. Mip-. doi:10.48550/arXiv.2111.12077 , urldate =. arxiv , file =:2111.12077 , primaryclass =

  76. [86]

    2018 , journal =

    Relational Inductive Biases, Deep Learning, and Graph Networks , author =. 2018 , journal =. 1806.01261 , archiveprefix =

  77. [87]

    2021 , journal =

    Paint by Word , author =. 2021 , journal =. 2103.10951 , archiveprefix =

  78. [88]

    doi:10.48550/arXiv.2207.13751 , urldate =

    Bautista, Miguel Angel and Guo, Pengsheng and Abnar, Samira and Talbott, Walter and Toshev, Alexander and Chen, Zhuoyuan and Dinh, Laurent and Zhai, Shuangfei and Goh, Hanlin and Ulbricht, Daniel and Dehghan, Afshin and Susskind, Josh , year =. doi:10.48550/arXiv.2207.13751 , ...

  79. [89]

    , year =

    Bello, Irwan and Zoph, Barret and Vaswani, Ashish and Shlens, Jonathon and Le, Quoc V. , year =. Attention Augmented Convolutional Networks , booktitle =

  80. [90]

    and Cohan, Arman , year =

    Beltagy, Iz and Peters, Matthew E. and Cohan, Arman , year =. Longformer:. arXiv preprint arXiv:2004.05150 , eprint =

  81. [91]

    2019 , keywords =

    Block:. 2019 , keywords =

  82. [92]

    Blattmann, Andreas and Rombach, Robin and Ling, Huan and Dockhorn, Tim and Kim, Seung Wook and Fidler, Sanja and Kreis, Karsten , year =. Align. Proceedings of the

  83. [93]

    Blattmann, Andreas and Milbich, Timo and Dorkenwald, Michael and Ommer, Bj. Ipoke:. Proceedings of the. 2021 , pages =

  84. [94]

    Understanding

    Blattmann, Andreas and Milbich, Timo and Dorkenwald, Michael and Ommer, Bjorn , year =. Understanding. Proceedings of the

  85. [95]

    Understanding

    Blattmann, Andreas and Milbich, Timo and Dorkenwald, Michael and Ommer, Bjorn , year =. Understanding

  86. [96]

    2003 , journal =

    Latent Dirichlet Allocation , author =. 2003 , journal =

  87. [97]

    1998 , journal =

    Creativity and Artificial Intelligence , author =. 1998 , journal =

  88. [98]

    2017 , month = jun, series =

    Bola. 2017 , month = jun, series =. doi:10.1007/978-3-319-58838-4_41 , urldate =

  89. [99]

    Bolya, Daniel and Fu, Cheng-Yang and Dai, Xiaoliang and Zhang, Peizhao and Feichtenhofer, Christoph and Hoffman, Judy , year =. Token. arXiv preprint arXiv:2210.09461 , eprint =

  90. [100]

    2021 , journal =

    Deep. 2021 , journal =. 2103.04922 , archiveprefix =

  91. [101]

    2014 , journal =

    From Machine Learning to Machine Reasoning , author =. 2014 , journal =

  92. [102]

    Boyd, Colin and Carr, Christopher , year =. Fair. Australasian

  93. [103]

    Digital Twin Technology in the Field Reclaims Offshore Resources , booktitle =

    Brewer, Thornton and Knight, Darrell and Noiray, Gautier and Naik, Harit , year =. Digital Twin Technology in the Field Reclaims Offshore Resources , booktitle =

  94. [104]

    Brock, Andrew and Donahue, Jeff and Simonyan, Karen , year =. Large. arXiv:1809.11096 [cs, stat] , eprint =

  95. [105]

    and Karras, Tero , year =

    Brooks, Tim and Hellsten, Janne and Aittala, Miika and Wang, Ting-Chun and Aila, Timo and Lehtinen, Jaakko and Liu, Ming-Yu and Efros, Alexei A. and Karras, Tero , year =. Generating. arXiv preprint arXiv:2206.03429 , eprint =

  96. [106]

    and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and

    Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D. and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and. Language. 2020 , journal =

  97. [107]

    , year =

    Burgert, Ryan and Ranasinghe, Kanchana and Li, Xiang and Ryoo, Michael S. , year =. Peekaboo:. doi:10.48550/arXiv.2211.13224 , urldate =. arxiv , file =:2211.13224 , primaryclass =

  98. [108]

    and Cudic, M

    Burt, R. and Cudic, M. and Principe, J. C. , year =. Fusing Attention with Visual Question Answering , booktitle =. doi:10.1109/IJCNN.2017.7965954 , abstract =

  99. [109]

    Contextvp:

    Byeon, Wonmin and Wang, Qin and Srivastava, Rupesh Kumar and Koumoutsakos, Petros , year =. Contextvp:. Proceedings of the

  100. [110]

    Architecture of the

    Cachin, Christian , year =. Architecture of the

  101. [111]

    Coco-Stuff:

    Caesar, Holger and Uijlings, Jasper and Ferrari, Vittorio , year =. Coco-Stuff:. Proceedings of the

  102. [112]

    doi:10.48550/arXiv.2211.12131 , urldate =

    Cai, Shengqu and Chan, Eric Ryan and Peng, Songyou and Shahbazi, Mohamad and Obukhov, Anton and Van Gool, Luc and Wetzstein, Gordon , year =. doi:10.48550/arXiv.2211.12131 , urldate =. arxiv , file =:2211.12131 , primaryclass =

  103. [113]

    Cao, Wei and Wang, Dong and Li, Jian and Zhou, Hao and Li, Lei and Li, Yitan , year =

  104. [114]

    Proceedings of the

    Cao, Ang and Rockwell, Chris and Johnson, Justin , year =. Proceedings of the

  105. [115]

    doi:10.48550/arXiv.2301.09632 , urldate =

    Cao, Ang and Johnson, Justin , year =. doi:10.48550/arXiv.2301.09632 , urldate =. arxiv , file =:2301.09632 , primaryclass =

  106. [116]

    Cao, Chenjie and Hong, Yuxin and Li, Xiang and Wang, Chengrong and Xu, Chengming and Xue, XiangYang and Fu, Yanwei , year =. The. arXiv preprint arXiv:2106.02514 , eprint =

  107. [117]

    Interpretable

    Cao, Qingxing and Liang, Xiaodan and Li, Bailin and Lin, Liang , year =. Interpretable. IEEE Transactions on Pattern Analysis and Machine Intelligence , keywords =

  108. [118]

    Cao, Liangfu and Gao, Lianli and Song, Jingkuan and Xu, Xing and Shen, Heng Tao , year =. Jointly. Databases. doi:10.1007/978-3-319-68155-9_19 , urldate =

  109. [119]

    Cao, Qingxing and Liang, Xiaodan and Li, Bailing and Li, Guanbin and Lin, Liang , year =. Visual

  110. [120]

    Carion, Nicolas and Massa, Francisco and Synnaeve, Gabriel and Usunier, Nicolas and Kirillov, Alexander and Zagoruyko, Sergey , year =. End-to-

  111. [121]

    Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset , shorttitle =

    Carreira, Joao and Zisserman, Andrew , year =. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset , shorttitle =. Proceedings of the

  112. [122]

    1997 , journal =

    Multitask Learning , author =. 1997 , journal =

  113. [123]

    Hierarchical

    Chandar, Sarath and Ahn, Sungjin and Larochelle, Hugo and Vincent, Pascal and Tesauro, Gerald and Bengio, Yoshua , year =. Hierarchical. arXiv preprint arXiv:1605.07427 , eprint =

  114. [124]

    Chandrasekaran, Arjun and Yadav, Deshraj and Chattopadhyay, Prithvijit and Prabhu, Viraj and Parikh, Devi , year =. It. arXiv:1704.00717 [cs] , eprint =

  115. [125]

    Textually

    Chandu, Khyathi Raghavi and Pyreddy, Mary Arpita and Felix, Matthieu and Joshi, Narendra Nath , year =. Textually. arXiv:1809.08697 [cs] , eprint =

  116. [126]

    and Lin, Connor Z

    Chan, Eric R. and Lin, Connor Z. and Chan, Matthew A. and Nagano, Koki and Pan, Boxiao and De Mello, Shalini and Gallo, Orazio and Guibas, Leonidas J. and Tremblay, Jonathan and Khamis, Sameh , year =. Efficient Geometry-Aware. Proceedings of the

  117. [127]

    , year =

    Chan, Caroline and Ginosar, Shiry and Zhou, Tinghui and Efros, Alexei A. , year =. Everybody. Proceedings of the

  118. [128]

    , year =

    Chang, Huiwen and Zhang, Han and Jiang, Lu and Liu, Ce and Freeman, William T. , year =. Maskgit:. Proceedings of the

  119. [129]

    Chang, Huiwen and Zhang, Han and Barber, Jarred and Maschinot, A. J. and Lezama, Jose and Jiang, Lu and Yang, Ming-Hsuan and Murphy, Kevin and Freeman, William T. and Rubinstein, Michael and Li, Yuanzhen and Krishnan, Dilip , year =. Muse:. doi:10.48550/arXiv.2301.00704 , urld...

  120. [130]

    Structure-

    Chang, Jianlong and Gu, Jie and Wang, Lingfeng and Meng, Gaofeng and Xiang, Shiming and Pan, Chunhong , year =. Structure-

  121. [131]

    Chao, Wei-Lun and Hu, Hexiang and Sha, Fei , year =. Being

  122. [132]

    Chao, Wei-Lun and Hu, Hexiang and Sha, Fei , year =. Cross-

  123. [133]

    Return of the Devil in the Details:

    Chatfield, Ken and Simonyan, Karen and Vedaldi, Andrea and Zisserman, Andrew , year =. Return of the Devil in the Details:. arXiv preprint arXiv:1405.3531 , eprint =

  124. [134]

    Evaluating Visual Conversational Agents via Cooperative Human-Ai Games , booktitle =

    Chattopadhyay, Prithvijit and Yadav, Deshraj and Prabhu, Viraj and Chandrasekaran, Arjun and Das, Abhishek and Lee, Stefan and Batra, Dhruv and Parikh, Devi , year =. Evaluating Visual Conversational Agents via Cooperative Human-Ai Games , booktitle =

  125. [135]

    Chen, Yunpeng and Kalantidis, Yannis and Li, Jianshu and Yan, Shuicheng and Feng, Jiashi , year =. A\^ 2-

  126. [136]

    arXiv:1511.05960 , eprint =

    Chen, Kan and Wang, Jiang and Chen, Liang-Chieh and Gao, Haoyuan and Xu, Wei and Nevatia, Ram , year =. arXiv:1511.05960 , eprint =

  127. [137]

    Character-

    Chen, Hong and Han, Rujun and Wu, Te-Lin and Nakayama, Hideki and Peng, Nanyun , year =. Character-. arXiv preprint arXiv:2210.08465 , eprint =

  128. [138]

    doi:10.48550/arXiv.2211.09788 , urldate =

    Chen, Shoufa and Sun, Peize and Song, Yibing and Luo, Ping , year =. doi:10.48550/arXiv.2211.09788 , urldate =. arxiv , keywords =:2211.09788 , primaryclass =

  129. [139]

    Generative Pretraining from Pixels , booktitle =

    Chen, Mark and Radford, Alec and Child, Rewon and Wu, Jeffrey and Jun, Heewoo and Luan, David and Sutskever, Ilya , year =. Generative Pretraining from Pixels , booktitle =

  130. [140]

    2016 , journal =

    Long Short-Term Memory-Networks for Machine Reading , author =. 2016 , journal =. arxiv , file =:1601.06733 , urldate =

  131. [141]

    2023 , journal =

    Segment and Track Anything , author =. 2023 , journal =. 2305.06558 , archiveprefix =

  132. [142]

    , year =

    Chen, Xinlei and Lawrence Zitnick, C. , year =. Mind's Eye:. Proceedings of the

  133. [143]

    Chen, Tsai-Shien and Lin, Chieh Hubert and Tseng, Hung-Yu and Lin, Tsung-Yi and Yang, Ming-Hsuan , year =. Motion-. arXiv preprint arXiv:2304.14404 , eprint =

  134. [144]

    2019 , journal =

    Residual Flows for Invertible Generative Modeling , author =. 2019 , journal =. 1906.02735 , archiveprefix =

  135. [145]

    and Li, Q

    Chen, L. and Li, Q. and Wang, H. and Long, Y. , year =. Static. doi:10.1109/BigComp.2018.00087 , abstract =

  136. [146]

    Teaching

    Chen, Xinyun and Lin, Maxwell and Sch. Teaching. 2023 , month = apr, number =. doi:10.48550/arXiv.2304.05128 , urldate =. arxiv , file =:2304.05128 , primaryclass =

  137. [147]

    doi:10.48550/arXiv.2102.10407 , urldate =

    Chen, Jun and Guo, Han and Yi, Kai and Li, Boyang and Elhoseiny, Mohamed , year =. doi:10.48550/arXiv.2102.10407 , urldate =. arxiv , file =:2102.10407 , primaryclass =

  138. [148]

    Recurrent

    Chiappa, Silvia and Racaniere, S. Recurrent. International. 2016 , month = nov, urldate =

  139. [149]

    Automatic

    Chi, Peggy and Sun, Zheng and Panovich, Katrina and Essa, Irfan , year =. Automatic. Proceedings of the 33rd

  140. [150]

    2020 , journal =

    Very Deep Vaes Generalize Autoregressive Models and Can Outperform Them on Images , author =. 2020 , journal =. 2011.10650 , archiveprefix =

  141. [151]

    doi:10.48550/arXiv.2108.02938 , urldate =

    Choi, Jooyoung and Kim, Sungwon and Jeong, Yonghyun and Gwon, Youngjune and Yoon, Sungroh , year =. doi:10.48550/arXiv.2108.02938 , urldate =. arxiv , file =:2108.02938 , primaryclass =

  142. [152]

    and Lee, W

    Cho, S. and Lee, W. H. and Kim, J. H. , year =. Implementation of Human-Robot. doi:10.1109/SMC.2017.8122654 , abstract =

  143. [153]

    Learning Phrase Representations Using

    Cho, Kyunghyun and Van Merri. Learning Phrase Representations Using. 2014 , journal =. arxiv , file =:1406.1078 , urldate =

  144. [154]

    Cho, Kyunghyun and. On the. 2014 , journal =. arxiv , file =:1409.1259 , urldate =

  145. [155]

    Choromanski, Krzysztof and Likhosherstov, Valerii and Dohan, David and Song, Xingyou and Davis, Jared and Sarlos, Tamas and Belanger, David and Colwell, Lucy and Weller, Adrian , year =. Masked. arXiv preprint arXiv:2006.03555 , eprint =

  146. [156]

    2022 , month = oct, number =

    Chowdhery, Aakanksha and Narang, Sharan and Devlin, Jacob and Bosma, Maarten and Mishra, Gaurav and Roberts, Adam and Barham, Paul and Chung, Hyung Won and Sutton, Charles and Gehrmann, Sebastian and Schuh, Parker and Shi, Kensen and Tsvyashchenko, Sasha and Maynez, Joshua and...

  147. [157]

    and Nguyen, K

    Chowdhury, I. and Nguyen, K. and Fookes, C. and Sridharan, S. , year =. A Cascaded Long Short-Term Memory (. doi:10.1109/ICIP.2017.8296600 , abstract =

  148. [158]

    Cho, Jaemin and Lu, Jiasen and Schwenk, Dustin and Hajishirzi, Hannaneh and Kembhavi, Aniruddha , year =. X-. Proceedings of the 2020. doi:10.18653/v1/2020.emnlp-main.707 , urldate =

  149. [159]

    Using Syntax to Ground Referring Expressions in Natural Images , booktitle =

    Cirik, Volkan and. Using Syntax to Ground Referring Expressions in Natural Images , booktitle =. 2018 , file =

  150. [160]

    2019 , journal =

    Adversarial Video Generation on Complex Datasets , author =. 2019 , journal =. 1907.06571 , archiveprefix =

  151. [161]

    , year =

    Clark, Kevin and Khandelwal, Urvashi and Levy, Omer and Manning, Christopher D. , year =. What Does Bert Look at? An Analysis of Bert's Attention , shorttitle =. arXiv preprint arXiv:1906.04341 , eprint =

  152. [162]

    2019 , journal =

    Visualizing and Measuring the Geometry of Bert , author =. 2019 , journal =. 1906.02715 , archiveprefix =

  153. [163]

    2019 , journal =

    On the Relationship between Self-Attention and Convolutional Layers , author =. 2019 , journal =. 1911.03584 , archiveprefix =

  154. [164]

    2019 , journal =

    Adaptively Sparse Transformers , author =. 2019 , journal =. 1909.00015 , archiveprefix =

  155. [165]

    2017 , journal =

    Unsupervised Learning from Video to Detect Foreground Objects in Single Images , author =. 2017 , journal =. arxiv , file =:1703.10901 , urldate =

  156. [166]

    2018 , month = may, journal =

    A Flexible Testing Environment for Visual Question Answering with Performance Evaluation , author =. 2018 , month = may, journal =. doi:10.1016/j.neucom.2018.02.065 , urldate =

  157. [167]

    Fine-Tune

    Cui, Baiyun and Li, Yingming and Chen, Ming and Zhang, Zhongfei , year =. Fine-Tune. Proceedings of the 2019

  158. [168]

    Dai, Zihang and Li, Lei and Xu, Wei , year =. Cfo:. arXiv preprint arXiv:1606.01994 , eprint =

  159. [169]

    Contrastive Learning for Image Captioning , booktitle =

    Dai, Bo and Lin, Dahua , year =. Contrastive Learning for Image Captioning , booktitle =

  160. [170]

    , year =

    Dai, Zihang and Lai, Guokun and Yang, Yiming and Le, Quoc V. , year =. Funnel-. arXiv preprint arXiv:2006.03236 , eprint =

  161. [171]

    doi:10.48550/arXiv.2305.06500 , urldate =

    Dai, Wenliang and Li, Junnan and Li, Dongxu and Tiong, Anthony Meng Huat and Zhao, Junqi and Wang, Weisheng and Li, Boyang and Fung, Pascale and Hoi, Steven , year =. doi:10.48550/arXiv.2305.06500 , urldate =. arxiv , file =:2305.06500 , primaryclass =

  162. [172]

    and Carbonell, Jaime and Le, Quoc V

    Dai, Zihang and Yang, Zhilin and Yang, Yiming and Cohen, William W. and Carbonell, Jaime and Le, Quoc V. and Salakhutdinov, Ruslan , year =. Transformer-Xl:. arXiv preprint arXiv:1901.02860 , eprint =

  163. [173]

    Proceedings of the 5th

    Dalvi, Bhavana and Bhakthavatsalam, Sumithra and Clark, Chris and Clark, Peter and Etzioni, Oren and Fader, Anthony and Groeneveld, Dirk , year =. Proceedings of the 5th

  164. [174]

    Scaling Egocentric Vision:

    Damen, Dima and Doughty, Hazel and Maria Farinella, Giovanni and Fidler, Sanja and Furnari, Antonino and Kazakos, Evangelos and Moltisanti, Davide and Munro, Jonathan and Perrett, Toby and Price, Will , year =. Scaling Egocentric Vision:. Proceedings of the

  165. [175]

    Das, Abhishek and Agrawal, Harsh and Zitnick, Larry and Parikh, Devi and Batra, Dhruv , year =. Human. Computer Vision and Image Understanding , series =. doi:10.1016/j.cviu.2017.10.001 , urldate =

  166. [176]

    and Corso, Jason J

    Das, Pradipto and Xu, Chenliang and Doell, Richard F. and Corso, Jason J. , year =. A Thousand Frames in Just a Few Words:. Proceedings of the

  167. [177]

    Visual Dialog , booktitle =

    Das, Abhishek and Kottur, Satwik and Gupta, Khushi and Singh, Avi and Yadav, Deshraj and Moura, Jos. Visual Dialog , booktitle =. 2017 , pages =

  168. [178]

    1990 , journal =

    Indexing by Latent Semantic Analysis , author =. 1990 , journal =

  169. [179]

    Companion

    Question. Companion. 2018 , series =. doi:10.1145/3184558.3192318 , urldate =

  170. [180]

    Clausie: Clause-Based Open Information Extraction , shorttitle =

    Del Corro, Luciano and Gemulla, Rainer , year =. Clausie: Clause-Based Open Information Extraction , shorttitle =. Proceedings of the 22nd International Conference on

  171. [181]

    Imagenet:

    Deng, Jia and Dong, Wei and Socher, Richard and Li, Li-Jia and Li, Kai and. Imagenet:. 2009. 2009 , pages =

  172. [182]

    Deng, Kangle and Fei, Tianyi and Huang, Xin and Peng, Yuxin , year =

  173. [183]

    Latent Alignment and Variational Attention , booktitle =

    Deng, Yuntian and Kim, Yoon and Chiu, Justin and Guo, Demi and Rush, Alexander , year =. Latent Alignment and Variational Attention , booktitle =

  174. [184]

    and Yan, Xinchen and Zhou, Yin and Guibas, Leonidas and Anguelov, Dragomir , year =

    Deng, Congyue and Jiang, Chiyu "Max'' and Qi, Charles R. and Yan, Xinchen and Zhou, Yin and Guibas, Leonidas and Anguelov, Dragomir , year =. doi:10.48550/arXiv.2212.03267 , urldate =. arxiv , file =:2212.03267 , primaryclass =

  175. [185]

    Unsupervised Object Region Proposals for

    Deng, Zhuo and Todorovic, Sinisa and Latecki, Longin Jan , year =. Unsupervised Object Region Proposals for. Computer Vision and Image Understanding , volume =

  176. [186]

    Visual Grounding via Accumulated Attention , booktitle =

    Deng, Chaorui and Wu, Qi and Wu, Qingyao and Hu, Fuyuan and Lyu, Fan and Tan, Mingkui , year =. Visual Grounding via Accumulated Attention , booktitle =

  177. [187]

    Stochastic Video Generation with a Learned Prior , booktitle =

    Denton, Emily and Fergus, Rob , year =. Stochastic Video Generation with a Learned Prior , booktitle =

  178. [188]

    2017 , journal =

    Unsupervised Learning of Disentangled Representations from Video , author =. 2017 , journal =. 1705.10915 , archiveprefix =

  179. [189]

    arXiv preprint arXiv:2006.06666 , eprint =

    Desai, Karan and Johnson, Justin , year =. arXiv preprint arXiv:2006.06666 , eprint =

  180. [190]

    Desta, M. T. and Chen, L. and Kornuta, T. , year =. Object-. doi:10.1109/WACV.2018.00201 , abstract =

  181. [191]

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , year =

  182. [192]

    Guesswhat?! Visual Object Discovery through Multi-Modal Dialogue , booktitle =

    De Vries, Harm and Strub, Florian and Chandar, Sarath and Pietquin, Olivier and Larochelle, Hugo and Courville, Aaron , year =. Guesswhat?! Visual Object Discovery through Multi-Modal Dialogue , booktitle =

  183. [193]

    Modulating Early Visual Processing by Language , booktitle =

    De Vries, Harm and Strub, Florian and Mary, J. Modulating Early Visual Processing by Language , booktitle =. 2017 , pages =

  184. [194]

    Diffusion Models Beat Gans on Image Synthesis , booktitle =

    Dhariwal, Prafulla and Nichol, Alexander , year =. Diffusion Models Beat Gans on Image Synthesis , booktitle =

  185. [195]

    and Salakhutdinov, Ruslan , year =

    Dhingra, Bhuwan and Yang, Zhilin and Cohen, William W. and Salakhutdinov, Ruslan , year =. Linguistic. arXiv preprint arXiv:1703.02620 , eprint =

  186. [196]

    arXiv preprint arXiv:2204.14217 , eprint =

    Ding, Ming and Zheng, Wendi and Hong, Wenyi and Tang, Jie , year =. arXiv preprint arXiv:2204.14217 , eprint =

  187. [197]

    Cogview:

    Ding, Ming and Yang, Zhuoyi and Hong, Wenyi and Zheng, Wendi and Zhou, Chang and Yin, Da and Lin, Junyang and Zou, Xu and Shao, Zhou and Yang, Hongxia , year =. Cogview:. Advances in

  188. [198]

    2016 , journal =

    Density Estimation Using Real Nvp , author =. 2016 , journal =. 1605.08803 , archiveprefix =

  189. [199]

    , year =

    Do, Tuong and Do, Thanh-Toan and Tran, Huy and Tjiputra, Erman and Tran, Quang D. , year =. Compact Trilinear Interaction for Visual Question Answering , booktitle =

  190. [200]

    Long-Term Recurrent Convolutional Networks for Visual Recognition and Description , booktitle =

    Donahue, Jeffrey and Anne Hendricks, Lisa and Guadarrama, Sergio and Rohrbach, Marcus and Venugopalan, Subhashini and Saenko, Kate and Darrell, Trevor , year =. Long-Term Recurrent Convolutional Networks for Visual Recognition and Description , booktitle =

  191. [201]

    Dong, Li and Yang, Nan and Wang, Wenhui and Wei, Furu and Liu, Xiaodong and Wang, Yu and Gao, Jianfeng and Zhou, Ming and Hon, Hsiao-Wuen , year =. Unified

  192. [202]

    and Ommer, Bjorn , year =

    Dorkenwald, Michael and Milbich, Timo and Blattmann, Andreas and Rombach, Robin and Derpanis, Konstantinos G. and Ommer, Bjorn , year =. Stochastic

  193. [203]

    Dosovitskiy, Alexey and Beyer, Lucas and Kolesnikov, Alexander and Weissenborn, Dirk and Zhai, Xiaohua and Unterthiner, Thomas and Dehghani, Mostafa and Minderer, Matthias and Heigold, Georg and Gelly, Sylvain , year =. An. arXiv preprint arXiv:2010.11929 , eprint =

  194. [204]

    Driess, Danny and Xia, Fei and Sajjadi, Mehdi S. M. and Lynch, Corey and Chowdhery, Aakanksha and Ichter, Brian and Wahid, Ayzaan and Tompson, Jonathan and Vuong, Quan and Yu, Tianhe and Huang, Wenlong and Chebotar, Yevgen and Sermanet, Pierre and Duckworth, Daniel and Levine,...

  195. [205]

    1973 , publisher =

    Pattern Classification and Scene Analysis , author =. 1973 , publisher =

  196. [206]

    Scam! Transferring Humans between Images with Semantic Cross Attention Modulation , booktitle =

    Dufour, Nicolas and Picard, David and Kalogeiton, Vicky , year =. Scam! Transferring Humans between Images with Semantic Cross Attention Modulation , booktitle =

  197. [207]

    2019 , journal =

    Implicit Generation and Generalization in Energy-Based Models , author =. 2019 , journal =. 1903.08689 , archiveprefix =

  198. [208]

    , year =

    Duke, Brendan and Taylor, Graham W. , year =. Generalized. arXiv:1803.09374 , eprint =

  199. [209]

    2018 , journal =

    Feature-Wise Transformations , author =. 2018 , journal =

  200. [210]

    and Levine, Sergey , year =

    Ebert, Frederik and Finn, Chelsea and Lee, Alex X. and Levine, Sergey , year =. Self-

  201. [211]

    Conductor , urldate =

    Educational. Conductor , urldate =

  202. [212]

    arXiv preprint arXiv:2112.05253 , eprint =

    Eichenberg, Constantin and Black, Sidney and Weinbach, Samuel and Parcalabescu, Letitia and Frank, Anette , year =. arXiv preprint arXiv:2112.05253 , eprint =

  203. [213]

    and Lippman, Andrew , year =

    Ekblaw, Ariel and Azaria, Asaph and Halamka, John D. and Lippman, Andrew , year =. A

  204. [214]

    1990 , journal =

    Finding Structure in Time , author =. 1990 , journal =

  205. [215]

    and Krishnan, Dilip and Mobahi, Hossein and Regan, Kevin and Bengio, Samy , year =

    Elsayed, Gamaleldin F. and Krishnan, Dilip and Mobahi, Hossein and Regan, Kevin and Bengio, Samy , year =. Large

  206. [216]

    , year =

    Elton, Daniel C. , year =. Self-Explaining. Artificial. doi:10.1007/978-3-030-52152-3_10 , urldate =

  207. [217]

    Endo, Yuki , year =. User-. Computer

  208. [218]

    Advances in

    Esser, Patrick and Rombach, Robin and Blattmann, Andreas and Ommer, Bjorn , year =. Advances in

  209. [219]

    2023 , journal =

    Structure and Content-Guided Video Synthesis with Diffusion Models , author =. 2023 , journal =. 2302.03011 , archiveprefix =

  210. [220]

    Taming Transformers for High-Resolution Image Synthesis , booktitle =

    Esser, Patrick and Rombach, Robin and Ommer, Bjorn , year =. Taming Transformers for High-Resolution Image Synthesis , booktitle =

  211. [221]

    Bitcoin-

    Eyal, Ittay and Gencer, Adem Efe and Sirer, Emin G. Bitcoin-. 13th. 2016 , pages =

  212. [222]

    2019 , month = sep, urldate =

  213. [223]

    Fan, Haoqi and Zhou, Jiatong , year =. Stacked

  214. [224]

    and Khan, Salman , year =

    Farazi, Moshiur R. and Khan, Salman , year =. Reciprocal. arXiv preprint arXiv:1805.04247 , eprint =

  215. [225]

    European Conference on Computer Vision , author =

    Every Picture Tells a Story:. European Conference on Computer Vision , author =. 2010 , pages =

  216. [226]

    Farha, Yazan Abu and Gall, Jurgen , year =. Ms-Tcn:. Proceedings of the

  217. [227]

    Proceedings of the 14th

    Farseev, Aleksandr and Yang, Qi and Filchenkov, Andrey and Lepikhin, Kirill and. Proceedings of the 14th. 2021 , pages =

  218. [228]

    , year =

    Fathi, Alireza and Li, Yin and Rehg, James M. , year =. Learning to Recognize Daily Actions Using Gaze , booktitle =

  219. [229]

    , year =

    Fathi, Alireza and Ren, Xiaofeng and Rehg, James M. , year =. Learning to Recognize Objects in Egocentric Activities , booktitle =. doi:10.1109/CVPR.2011.5995444 , abstract =

  220. [230]

    , year =

    Fathi, Alireza and Rehg, James M. , year =. Modeling Actions through State Changes , booktitle =

  221. [231]

    , year =

    Fathi, Alireza and Farhadi, Ali and Rehg, James M. , year =. Understanding Egocentric Activities , booktitle =

  222. [232]

    Slowfast Networks for Video Recognition , booktitle =

    Feichtenhofer, Christoph and Fan, Haoqi and Malik, Jitendra and He, Kaiming , year =. Slowfast Networks for Video Recognition , booktitle =

  223. [233]

    , year =

    Feigenbaum, Edward A. , year =. The Art of Artificial Intelligence. 1

  224. [234]

    Progressive

    Fei, Zhengcong and Fan, Mingyuan and Zhu, Li and Huang, Junshi and Wei, Xiaoming and Wei, Xiaolin , year =. Progressive. doi:10.48550/arXiv.2210.02291 , urldate =. arxiv , file =:2210.02291 , primaryclass =

  225. [235]

    Unsupervised Learning for Physical Interaction through Video Prediction , booktitle =

    Finn, Chelsea and Goodfellow, Ian and Levine, Sergey , year =. Unsupervised Learning for Physical Interaction through Video Prediction , booktitle =

  226. [236]

    Stochastic Latent Residual Video Prediction , booktitle =

    Franceschi, Jean-Yves and Delasalles, Edouard and Chen, Micka. Stochastic Latent Residual Video Prediction , booktitle =. 2020 , pages =

  227. [238]

    doi:10.48550/arXiv.2302.01133 , urldate =

    Fridman, Rafail and Abecasis, Amit and Kasten, Yoni and Dekel, Tali , year =. doi:10.48550/arXiv.2302.01133 , urldate =

  228. [239]

    Proceedings of the

    Plenoxels:. Proceedings of the. 2022 , pages =

  229. [240]

    Adversarial Text-to-Image Synthesis:

    Frolov, Stanislav and Hinz, Tobias and Raue, Federico and Hees, J. Adversarial Text-to-Image Synthesis:. 2021 , journal =. 2101.09983 , archiveprefix =

  230. [241]

    Tilegan: Synthesis of Large-Scale Non-Homogeneous Textures , shorttitle =

    Fr. Tilegan: Synthesis of Large-Scale Non-Homogeneous Textures , shorttitle =. 2019 , journal =

  231. [242]

    Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding , booktitle =

    Fukui, Akira and Park, Dong Huk and Yang, Daylen and Rohrbach, Anna and Darrell, Trevor and Rohrbach, Marcus , year =. Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding , booktitle =

  232. [243]

    Multi-Modal

    Gabeur, Valentin and Sun, Chen and Alahari, Karteek and Schmid, Cordelia , year =. Multi-Modal. European

  233. [244]

    Make-a-Scene:

    Gafni, Oran and Polyak, Adam and Ashual, Oron and Sheynin, Shelly and Parikh, Devi and Taigman, Yaniv , year =. Make-a-Scene:. arXiv preprint arXiv:2203.13131 , eprint =

  234. [245]

    1994 , journal =

    A New Algorithm for Data Compression , author =. 1994 , journal =

  235. [246]

    Gandikota, Rohit and Materzynska, Joanna and. Erasing. 2023 , month = jun, number =. doi:10.48550/arXiv.2303.07345 , urldate =. arxiv , file =:2303.07345 , primaryclass =

  236. [247]

    Hyperbolic

    Ganea, Octavian-Eugen and B. Hyperbolic. 2018 , file =

  237. [248]

    and Russakovsky, O

    Ganju, S. and Russakovsky, O. and Gupta, A. , year =. What's in a. doi:10.1109/CVPR.2017.680 , abstract =

  238. [249]

    Stylenet:

    Gan, Chuang and Gan, Zhe and He, Xiaodong and Gao, Jianfeng and Deng, Li , year =. Stylenet:. Proceedings of the

  239. [250]

    arxiv , file =:1708.04686 , abstract =

    Gan, Chuang and Li, Yandong and Li, Haoxiang and Sun, Chen and Gong, Boqing , year =. arxiv , file =:1708.04686 , abstract =

  240. [251]

    Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question , shorttitle =

    Gao, Haoyuan and Mao, Junhua and Zhou, Jie and Huang, Zhiheng and Wang, Lei and Xu, Wei , year =. Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question , shorttitle =

  241. [252]

    Gao, Jun and Shen, Tianchang and Wang, Zian and Chen, Wenzheng and Yin, Kangxue and Li, Daiqing and Litany, Or and Gojcic, Zan and Fidler, Sanja , year =. Get3d:. Advances

  242. [253]

    doi:10.48550/arXiv.2209.11163 , urldate =

    Gao, Jun and Shen, Tianchang and Wang, Zian and Chen, Wenzheng and Yin, Kangxue and Li, Daiqing and Litany, Or and Gojcic, Zan and Fidler, Sanja , year =. doi:10.48550/arXiv.2209.11163 , urldate =. arxiv , file =:2209.11163 , primaryclass =

  243. [254]

    doi:10.48550/arXiv.2304.15010 , urldate =

    Gao, Peng and Han, Jiaming and Zhang, Renrui and Lin, Ziyi and Geng, Shijie and Zhou, Aojun and Zhang, Wei and Lu, Pan and He, Conghui and Yue, Xiangyu and Li, Hongsheng and Qiao, Yu , year =. doi:10.48550/arXiv.2304.15010 , urldate =. arxiv , file =:2304.15010 , primaryclass =

  244. [255]

    and Kiayias, Aggelos and Leonardos, Nikos and Panagiotakos, Giorgos , year =

    Garay, Juan A. and Kiayias, Aggelos and Leonardos, Nikos and Panagiotakos, Giorgos , year =. Bootstrapping the

  245. [256]

    现代信息科技 , volume =

    葛, 梦颖 and 孙, 宝山 , year =. 现代信息科技 , volume =

  246. [257]

    2022 , journal =

    Long Video Generation with Time-Agnostic Vqgan and Time-Sensitive Transformer , author =. 2022 , journal =. 2204.03638 , archiveprefix =

  247. [258]

    Planting a

    Ge, Yuying and Ge, Yixiao and Zeng, Ziyun and Wang, Xintao and Shan, Ying , year =. Planting a. doi:10.48550/arXiv.2307.08041 , urldate =. arxiv , file =:2307.08041 , primaryclass =

  248. [259]

    and Torr, Philip HS and Dokania, Puneet K

    Ghosh, Arnab and Kulharia, Viveka and Namboodiri, Vinay P. and Torr, Philip HS and Dokania, Puneet K. , year =. Multi-Agent Diverse Generative Adversarial Networks , booktitle =

  249. [260]

    Stacked Spatio-Temporal Graph Convolutional Networks for Action Segmentation , booktitle =

    Ghosh, Pallabi and Yao, Yi and Davis, Larry and Divakaran, Ajay , year =. Stacked Spatio-Temporal Graph Convolutional Networks for Action Segmentation , booktitle =

  250. [261]

    Fast R-Cnn , booktitle =

    Girshick, Ross , year =. Fast R-Cnn , booktitle =

  251. [262]

    Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation , booktitle =

    Girshick, Ross and Donahue, Jeff and Darrell, Trevor and Malik, Jitendra , year =. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation , booktitle =

  252. [263]

    A Game-Theoretic Approach to Generating Spatial Descriptions , booktitle =

    Golland, Dave and Liang, Percy and Klein, Dan , year =. A Game-Theoretic Approach to Generating Spatial Descriptions , booktitle =

  253. [264]

    Improving Image-Sentence Embeddings Using Large Weakly Annotated Photo Collections , booktitle =

    Gong, Yunchao and Wang, Liwei and Hodosh, Micah and Hockenmaier, Julia and Lazebnik, Svetlana , year =. Improving Image-Sentence Embeddings Using Large Weakly Annotated Photo Collections , booktitle =

  254. [265]

    Generative Adversarial Nets , booktitle =

    Goodfellow, Ian and. Generative Adversarial Nets , booktitle =. 2014 , pages =

  255. [266]

    arxiv , file =:1712.03316 , urldate =

    Gordon, Daniel and Kembhavi, Aniruddha and Rastegari, Mohammad and Redmon, Joseph and Fox, Dieter and Farhadi, Ali , year =. arxiv , file =:1712.03316 , urldate =

  256. [267]

    Making the

    Goyal, Yash and Khot, Tejas and. Making the. 2017 , volume =

  257. [268]

    Goyal, Yash and Mohapatra, Akrit and Parikh, Devi and Batra, Dhruv , year =. Towards. arXiv:1608.08974 , eprint =

  258. [269]

    arXiv preprint arXiv:2110.07058 , eprint =

    Grauman, Kristen and Westbury, Andrew and Byrne, Eugene and Chavis, Zachary and Furnari, Antonino and Girdhar, Rohit and Hamburger, Jackson and Jiang, Hao and Liu, Miao and Liu, Xingyu , year =. arXiv preprint arXiv:2110.07058 , eprint =

  259. [270]

    and Kim, S

    Gu, G. and Kim, S. T. and Ro, Y. M. , year =. Adaptive Attention Fusion Network for Visual Question Answering , booktitle =. doi:10.1109/ICME.2017.8019540 , abstract =

  260. [271]

    Gulcehre, Caglar and Chandar, Sarath and Cho, Kyunghyun and Bengio, Yoshua , year =. Dynamic. arXiv preprint arXiv:1607.00036 , eprint =

  261. [272]

    Hyperbolic

    Gulcehre, Caglar and Denil, Misha and Malinowski, Mateusz and Razavi, Ali and Pascanu, Razvan and Hermann, Karl Moritz and Battaglia, Peter and Bapst, Victor and Raposo, David and Santoro, Adam , year =. Hyperbolic

  262. [273]

    Guo, Zonghui and Guo, Dongsheng and Zheng, Haiyong and Gu, Zhaorui and Zheng, Bing and Dong, Junyu , year =. Image. Proceedings of the

  263. [274]

    Proceedings of the

    Guo, Longteng and Liu, Jing and Yao, Peng and Li, Jiangwei and Lu, Hanqing , year =. Proceedings of the

  264. [275]

    2019 , journal =

    Star-Transformer , author =. 2019 , journal =. 1902.09113 , archiveprefix =

  265. [276]

    and Singh, Saurabh and Hoiem, Derek , year =

    Gupta, Tanmay and Shih, Kevin J. and Singh, Saurabh and Hoiem, Derek , year =. Aligned

  266. [277]

    arXiv preprint arXiv:2006.03274 , eprint =

    Gupta, Ankit and Berant, Jonathan , year =. arXiv preprint arXiv:2006.03274 , eprint =

  267. [278]

    Imagine This! Scripts to Compositions to Videos , booktitle =

    Gupta, Tanmay and Schwenk, Dustin and Farhadi, Ali and Hoiem, Derek and Kembhavi, Aniruddha , year =. Imagine This! Scripts to Compositions to Videos , booktitle =

  268. [279]

    Maskvit:

    Gupta, Agrim and Tian, Stephen and Zhang, Yunzhi and Wu, Jiajun and. Maskvit:. 2022 , journal =. 2206.11894 , archiveprefix =

  269. [280]

    Perceptual Organization and Recognition of Indoor Scenes from

    Gupta, Saurabh and Arbelaez, Pablo and Malik, Jitendra , year =. Perceptual Organization and Recognition of Indoor Scenes from. Proceedings of the

  270. [281]

    Proceedings of the

    Gupta, Sonam and Keshari, Arti and Das, Sukhendu , year =. Proceedings of the

  271. [282]

    Survey of

    Gupta, Akshay Kumar , year =. Survey of. arXiv:1705.03865 [cs] , eprint =

  272. [283]

    Gurari, Danna and Grauman, Kristen , year =

  273. [284]

    and Guo, Anhong and Lin, Chi and Grauman, Kristen and Luo, Jiebo and Bigham, Jeffrey P

    Gurari, Danna and Li, Qing and Stangl, Abigale J. and Guo, Anhong and Lin, Chi and Grauman, Kristen and Luo, Jiebo and Bigham, Jeffrey P. , year =. CVPR , file =

  274. [285]

    Gururangan, Suchin and Marasovi. Don't. 2020 , journal =. 2004.10964 , archiveprefix =

  275. [286]

    Vector Quantized Diffusion Model for Text-to-Image Synthesis , booktitle =

    Gu, Shuyang and Chen, Dong and Bao, Jianmin and Wen, Fang and Zhang, Bo and Chen, Dongdong and Yuan, Lu and Guo, Baining , year =. Vector Quantized Diffusion Model for Text-to-Image Synthesis , booktitle =

  276. [287]

    Automated

    Ha, JungWoo and Kim, Kyung-Min and Zhang, Byoung-Tak , year =. Automated

  277. [288]

    Han, Yuxuan and Wang, Ruicheng and Yang, Jiaolong , year =. Single-. arXiv preprint arXiv:2205.11733 , eprint =

  278. [289]

    Controllable

    Hao, Zekun and Huang, Xun and Belongie, Serge , year =. Controllable. Proceedings of the

  279. [290]

    and Holynski, Aleksander and Kanazawa, Angjoo , year =

    Haque, Ayaan and Tancik, Matthew and Efros, Alexei A. and Holynski, Aleksander and Kanazawa, Angjoo , year =. Instruct-. doi:10.48550/arXiv.2303.12789 , urldate =. arxiv , file =:2303.12789 , primaryclass =

  280. [291]

    Flexible

    Harvey, William and Naderiparizi, Saeid and Masrani, Vaden and Weilbach, Christian and Wood, Frank , year =. Flexible. arXiv preprint arXiv:2205.11495 , eprint =

  281. [292]

    Harzig, Philipp and Eggert, Christian and Lienhart, Rainer , year =. Visual. doi:10.1145/3206025.3206054 , urldate =

  282. [293]

    Deep Residual Learning for Image Recognition , booktitle =

    He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian , year =. Deep Residual Learning for Image Recognition , booktitle =

  283. [294]

    He, Yingqing and Yang, Tianyu and Zhang, Yong and Shan, Ying and Chen, Qifeng , year =. Latent. doi:10.48550/arXiv.2211.13221 , urldate =

  284. [295]

    He, Kaiming and Chen, Xinlei and Xie, Saining and Li, Yanghao and Doll. Masked. Proceedings of the. 2022 , pages =

  285. [296]

    He, Kaiming and Gkioxari, Georgia and Dollar, Piotr and Girshick, Ross , year =. Mask. Proceedings of the

  286. [297]

    Momentum

    He, Kaiming and Fan, Haoqi and Wu, Yuxin and Xie, Saining and Girshick, Ross , year =. Momentum

  287. [298]

    and Malcolm, George L

    Henderson, John M. and Malcolm, George L. and Schandl, Charles , year =. Searching in the Dark:. Psychonomic bulletin & review , volume =

  288. [299]

    Uncertainty-

    Heo, Jay and Lee, Hae Beom and Kim, Saehoon and Lee, Juho and Kim, Kwang Joon and Yang, Eunho and Hwang, Sung Ju , year =. Uncertainty-

  289. [300]

    Revisiting

    He, Junxian and Gu, Jiatao and Shen, Jiajun and Ranzato, Marc'Aurelio , year =. Revisiting. International

Pith tools

Reviewed May 20, 2026 · model on record in the stance chip above.