High-fidelity 3D Generation from images
Audio Conditioned LipSync with Latent Diffusion Models
text-to-image