ALM2Vec learns unified audio embeddings from large audio-language models for text-audio retrieval, instruction-aware retrieval, and other tasks across domains.
Learning transferable visual models from natural language supervision
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
A geometry-aware 4D video generation model trained with cross-view pointmap alignment to produce spatio-temporally consistent future videos from novel viewpoints for robot manipulation.
SMPL-GPTexture uses text-to-image generation to produce dual-view human images, aligns them to SMPL meshes via 2D-to-3D recovery, projects colors to UV space, and applies diffusion inpainting to create full high-resolution textures aligned to user prompts.
citing papers explorer
-
ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models
ALM2Vec learns unified audio embeddings from large audio-language models for text-audio retrieval, instruction-aware retrieval, and other tasks across domains.
-
Geometry-aware 4D Video Generation for Robot Manipulation
A geometry-aware 4D video generation model trained with cross-view pointmap alignment to produce spatio-temporally consistent future videos from novel viewpoints for robot manipulation.
-
SMPL-GPTexture: Dual-View 3D Human Texture Estimation using Text-to-Image Generation Models
SMPL-GPTexture uses text-to-image generation to produce dual-view human images, aligns them to SMPL meshes via 2D-to-3D recovery, projects colors to UV space, and applies diffusion inpainting to create full high-resolution textures aligned to user prompts.