Diffusers documentation
Kandinsky 6
Kandinsky 6
Kandinsky 6 is a family of video generation models from Kandinsky Lab. The main model generates video and synchronized audio from text or a reference image with a single multimodal diffusion transformer: video and audio latents are denoised together through fused blocks that cross-attend between the two modalities, each conditioned on its own Qwen2.5-VL text branch and a CLIP pooled embedding. A separate super-resolution model upscales the generated video tile by tile in the latent space of a causal 3D K-VAE.
Check out the Kandinsky Lab organization on the Hub for the full set of official checkpoints, including flow-matching and distilled variants of both the base and super-resolution models.
Distilled checkpoints ship with the few-step PiflowScheduler and must be run with
guidance_scale=1.0.
Available models
| Model | Pipeline | Notes |
|---|---|---|
kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers | Kandinsky6TI2VAPipeline | Flow matching, guidance_scale=5.0, 50 steps |
kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers | Kandinsky6TI2VAPipeline | Distilled, guidance_scale=1.0, 10 steps |
kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers | Kandinsky6TI2VAPipeline | Flow matching, guidance_scale=5.0, 50 steps |
kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers | Kandinsky6TI2VAPipeline | Distilled, guidance_scale=1.0, 10 steps |
kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers | Kandinsky6TI2VAPipeline | Flow matching, guidance_scale=5.0, 50 steps |
kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers | Kandinsky6TI2VAPipeline | Flow matching, guidance_scale=5.0, 50 steps |
kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers | Kandinsky6SRPipeline | Flow matching super-resolution |
kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers | Kandinsky6SRPipeline | Distilled super-resolution, 2 steps |
Text/image-to-video-and-audio
import torch
from diffusers import Kandinsky6TI2VAPipeline
from diffusers.utils import encode_video
pipe = Kandinsky6TI2VAPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
output = pipe(
prompt="A cat and a dog baking a cake together in a kitchen.",
height=480,
width=864,
num_frames=121,
num_inference_steps=10,
guidance_scale=1.0,
)
encode_video(
output.frames[0],
fps=24,
output_path="output.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)Pass image= to condition the first frame on a reference image, sample_audio=False to generate video only, and expand_prompts=True to let the Qwen2.5-VL text encoder rewrite short prompts into detailed ones first.
Video super-resolution
Kandinsky6SRPipeline takes the frames produced by Kandinsky6TI2VAPipeline and upscales them by 2, 4, or 2.25 (a 1.125x bilinear pre-upscale followed by the 2x path). The video is split into overlapping tiles, every tile
is refined at one of the tile sizes the SR transformer was trained on, and the tiles are blended back with Hann
windows.
# required: lets inductor pick flex-attention tiles that fit the SR block mask
torch._inductor.config.max_autotune = True
sr_pipe = Kandinsky6SRPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", torch_dtype=torch.bfloat16
)
# The SR transformer always runs NABLA sparse attention on the `flex` backend. Compile it, otherwise flex falls
# back to an eager implementation that needs far more memory at video resolutions.
sr_pipe.enable_model_cpu_offload()
sr_pipe.transformer.set_attention_backend("flex")
sr_pipe.transformer.compile_repeated_blocks(fullgraph=True)
upscaled = sr_pipe(video=output.frames[0], resolution_scale=2.25, num_inference_steps=2).frames[0]Memory optimization
Refer to the Reduce memory usage guide for the general set of techniques. Both Kandinsky6TI2VAPipeline and Kandinsky6SRPipeline support model offloading (used above) and, for a smaller footprint at the cost of speed, sequential CPU offloading:
pipe.enable_sequential_cpu_offload()
Kandinsky6TI2VAPipeline’s video VAE also supports tiled decoding for high resolutions or long videos:
pipe.vae.enable_tiling()
Notes
heightandwidthmust be divisible by the video VAE’s spatial compression ratio times the transformer’s patch size —16with the default Kandinsky6TI2VAPipeline configuration (AutoencoderKLHunyuanVideoat a compression ratio of8,patch_size=(1, 2, 2)).480x864, used in the example above, satisfies this.- Kandinsky6SRPipeline’s input
videomust have1 + k * vae_scale_factor_temporalframes for some integerk(a temporal compression ratio of4with the default K-VAE configuration, so121frames works but120doesn’t) — trim or pad a video that doesn’t already satisfy this before upscaling it. expand_prompts=Truereuses the already-loaded Qwen2.5-VL text encoder for an extra generation pass before denoising, so it adds latency but no extra model weights.- Compile the repeated transformer blocks for faster repeated inference:
pipe.transformer.compile_repeated_blocks(fullgraph=True)
Kandinsky6TI2VAPipeline
class diffusers.Kandinsky6TI2VAPipeline
< source >( transformer: Kandinsky6Transformer3DModelvae: AutoencoderKLHunyuanVideotext_encoder: Qwen2_5_VLForConditionalGenerationtokenizer: Qwen2_5_VLProcessortext_encoder_2: CLIPTextModeltokenizer_2: CLIPTokenizerscheduler: diffusers.schedulers.scheduling_flow_match_euler_discrete.FlowMatchEulerDiscreteScheduler | diffusers.schedulers.scheduling_piflow.PiflowScheduleraudio_vae: diffusers.models.autoencoders.autoencoder_mmaudio.MMAudioVAE | None = Nonevocoder: diffusers.models.autoencoders.mmaudio_vocoder.MMAudioVocoder | None = None )
Parameters
- transformer (Kandinsky6Transformer3DModel) — Multimodal transformer that denoises the video and audio latents.
- vae (AutoencoderKLHunyuanVideo) — Video VAE used to encode the reference image and decode the generated video.
- text_encoder (Qwen2_5_VLForConditionalGeneration) — Qwen2.5-VL model providing the token-level text embeddings and, optionally, prompt expansion.
- tokenizer (Qwen2_5_VLProcessor) —
Processor of
text_encoder. - text_encoder_2 (CLIPTextModel) — CLIP text encoder providing the pooled text embedding.
- tokenizer_2 (CLIPTokenizer) —
Tokenizer of
text_encoder_2. - scheduler (FlowMatchEulerDiscreteScheduler or PiflowScheduler) —
Scheduler used with
transformerto denoise the latents. Distilled checkpoints ship with a PiflowScheduler and must be run withguidance_scale=1.0. - audio_vae (MMAudioVAE, optional) —
Audio VAE used to decode the generated audio latents into a mel spectrogram. Only needed when
sample_audio=True. - vocoder (MMAudioVocoder, optional) —
Vocoder used to turn the mel spectrogram
audio_vaedecodes into a waveform. Only needed whensample_audio=True.
Pipeline for text/image-to-video-and-audio generation with Kandinsky 6.
Video and audio latents are denoised together by a single multimodal transformer, conditioned on Qwen2.5-VL text tokens and a CLIP pooled embedding. An optional reference image conditions the first frame.
This model inherits from DiffusionPipeline. Check the superclass documentation for the generic methods implemented for all pipelines (downloading, saving, running on a particular device, etc.).
__call__
< source >( prompt: str | list[str] | None = Noneimage: typing.Union[PIL.Image.Image, numpy.ndarray, torch.Tensor, list[PIL.Image.Image], list[numpy.ndarray], list[torch.Tensor], NoneType] = Nonenegative_prompt: str | list[str] | None = Noneheight: int = 512width: int = 768num_frames: int = 121frame_rate: float = 24.0num_inference_steps: int = 50timesteps: list[int] | None = Nonesigmas: list[float] | None = Noneguidance_scale: float = 5.0num_videos_per_prompt: int = 1generator: typing.Union[torch.Generator, list[torch.Generator], NoneType] = Nonelatents: typing.Optional[torch.Tensor] = Noneaudio_latents: typing.Optional[torch.Tensor] = Noneprompt_embeds: typing.Optional[torch.Tensor] = Nonepooled_prompt_embeds: typing.Optional[torch.Tensor] = Nonenegative_prompt_embeds: typing.Optional[torch.Tensor] = Nonenegative_pooled_prompt_embeds: typing.Optional[torch.Tensor] = Nonesample_audio: bool = Trueexpand_prompts: bool = Falsemax_sequence_length: int = 1024output_type: str = 'pil'return_dict: bool = Truecallback_on_step_end: collections.abc.Callable[[int, int, dict], None] | None = Nonecallback_on_step_end_tensor_inputs: list = ['latents'] ) → Kandinsky6TI2VAPipelineOutput or tuple
Parameters
- prompt (
strorlist[str], optional) — The prompt or prompts to guide the generation. Required unlessprompt_embedsis given. - image (
PipelineImageInput, optional) — Reference image(s) conditioning the first frame (image-to-video-and-audio). - negative_prompt (
strorlist[str], optional) — The prompt or prompts not to guide the generation. Defaults to the Kandinsky 6 negative prompt. - height (
int, defaults to512) — Height of the generated video in pixels. - width (
int, defaults to768) — Width of the generated video in pixels. - num_frames (
int, defaults to121) — Number of generated frames. - frame_rate (
float, defaults to24.0) — Frame rate the video is generated at; sets the length of the synchronized audio. - num_inference_steps (
int, defaults to50) — The number of denoising steps. Use16with the distilled checkpoints. - timesteps (
list[int], optional) — Custom timesteps for schedulers that support them. - sigmas (
list[float], optional) — Custom sigmas for schedulers that support them. - guidance_scale (
float, defaults to5.0) — Classifier-free guidance scale. Must be1.0with a PiflowScheduler. - num_videos_per_prompt (
int, defaults to1) — The number of videos to generate per prompt. - generator (
torch.Generatororlist[torch.Generator], optional) — Generator(s) used for the initial noise and the reference image encoding. - latents (
torch.Tensor, optional) — Pre-generated video latents of shape(batch_size, channels, num_latent_frames, latent_height, latent_width). - audio_latents (
torch.Tensor, optional) — Pre-generated audio latents of shape(batch_size, channels, audio_length). - prompt_embeds (
torch.Tensor, optional) — Pre-generated Qwen2.5-VL text embeddings. - pooled_prompt_embeds (
torch.Tensor, optional) — Pre-generated CLIP pooled text embeddings. - negative_prompt_embeds (
torch.Tensor, optional) — Pre-generated negative Qwen2.5-VL text embeddings. - negative_pooled_prompt_embeds (
torch.Tensor, optional) — Pre-generated negative CLIP pooled text embeddings. - sample_audio (
bool, defaults toTrue) — Whether to generate synchronized audio. Requires the pipeline to have anaudio_vaeand avocoder. - expand_prompts (
bool, defaults toFalse) — Whether to rewrite the prompts with expand_prompts() before encoding. - max_sequence_length (
int, defaults to1024) — Maximum number of prompt tokens after the chat template. - output_type (
str, defaults to"pil") — The output format of the generated video:"pil","np","pt"or"latent". - return_dict (
bool, defaults toTrue) — Whether or not to return a Kandinsky6TI2VAPipelineOutput instead of a plain tuple. - callback_on_step_end (
Callable, optional) — A function called at the end of each denoising step withcallback_on_step_end(self, step, timestep, callback_kwargs). It may return a dict overriding the listed tensors. - callback_on_step_end_tensor_inputs (
list[str], defaults to["latents"]) — Tensor inputs passed tocallback_on_step_end; a subset of_callback_tensor_inputs.
Returns
Kandinsky6TI2VAPipelineOutput or tuple
The generated video and audio; a (frames, audio) tuple when return_dict=False.
The call function to the pipeline for generation.
Examples:
>>> import torch
>>> from diffusers import Kandinsky6TI2VAPipeline
>>> from diffusers.utils import encode_video
>>> pipe = Kandinsky6TI2VAPipeline.from_pretrained(
... "kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", torch_dtype=torch.bfloat16
... )
>>> pipe.enable_model_cpu_offload()
>>> output = pipe(
... prompt="A cat and a dog baking a cake together in a kitchen.",
... height=480,
... width=864,
... num_frames=121,
... num_inference_steps=16,
... guidance_scale=1.0,
... )
>>> encode_video(
... output.frames[0],
... fps=24,
... output_path="output.mp4",
... audio=output.audio[0][None],
... audio_sample_rate=pipe.audio_sample_rate,
... )encode_image
< source >( image: typing.Union[PIL.Image.Image, numpy.ndarray, torch.Tensor, list[PIL.Image.Image], list[numpy.ndarray], list[torch.Tensor]]height: intwidth: intdevice: devicedtype: dtypenum_videos_per_prompt: int = 1generator: typing.Optional[torch.Generator] = None )
Encodes the reference image(s) into first-frame latents of shape (batch_size, latent_height, latent_width, latent_channels), scaled by the VAE scaling_factor. PIL images are resized and center-cropped to height x width; tensors and arrays must already have that size. The latents are repeated num_videos_per_prompt times
along the batch dimension.
encode_prompt
< source >( prompt: str | list[str]negative_prompt: str | list[str] | None = Nonedo_classifier_free_guidance: bool = Truenum_videos_per_prompt: int = 1prompt_embeds: typing.Optional[torch.Tensor] = Nonepooled_prompt_embeds: typing.Optional[torch.Tensor] = Noneprompt_attention_mask: typing.Optional[torch.Tensor] = Nonenegative_prompt_embeds: typing.Optional[torch.Tensor] = Nonenegative_pooled_prompt_embeds: typing.Optional[torch.Tensor] = Nonenegative_prompt_attention_mask: typing.Optional[torch.Tensor] = Nonemax_sequence_length: int = 1024device: typing.Optional[torch.device] = Nonedtype: typing.Optional[torch.dtype] = None )
Parameters
- prompt (
strorlist[str]) — Prompt to be encoded. - negative_prompt (
strorlist[str], optional) — The prompt not to guide the generation. Ignored whendo_classifier_free_guidanceisFalse. - do_classifier_free_guidance (
bool, defaults toTrue) — Whether to also encode the negative prompt. - num_videos_per_prompt (
int, defaults to1) — Number of videos generated per prompt; the embeddings are repeated accordingly. - prompt_embeds (
torch.Tensor, optional) — Pre-generated Qwen2.5-VL text embeddings. Skips encodingprompt. - pooled_prompt_embeds (
torch.Tensor, optional) — Pre-generated CLIP pooled text embeddings. Must be given together withprompt_embeds. - prompt_attention_mask (
torch.Tensor, optional) — Boolean padding mask ofprompt_embeds. - negative_prompt_embeds (
torch.Tensor, optional) — Pre-generated negative Qwen2.5-VL text embeddings. - negative_pooled_prompt_embeds (
torch.Tensor, optional) — Pre-generated negative CLIP pooled text embeddings. - negative_prompt_attention_mask (
torch.Tensor, optional) — Boolean padding mask ofnegative_prompt_embeds. - max_sequence_length (
int, defaults to1024) — Maximum number of prompt tokens after the chat template. - device (
torch.device, optional) — Device to run the text encoders on. - dtype (
torch.dtype, optional) — Dtype of the returned embeddings.
Encodes the prompt into text encoder hidden states.
expand_prompts
< source >( prompt: str | list[str]tokenizertext_encoderdevice: deviceimage: PIL.Image.Image | list[PIL.Image.Image] | None = Nonemax_sequence_length: int = 1024generator: typing.Union[torch.Generator, list[torch.Generator], NoneType] = None ) → str or list[str]
Parameters
- prompt (
strorlist[str]) — Prompt or prompts to expand. - tokenizer —
The Qwen2.5-VL processor, e.g.
pipe.tokenizer. - text_encoder —
The Qwen2.5-VL model, e.g.
pipe.text_encoder. - device (
torch.device) — Device to run the text encoder on. - image (
PIL.Image.Imageorlist[PIL.Image.Image], optional) — Reference image(s) of an image-to-video call. - max_sequence_length (
int, defaults to1024) — Maximum number of generated tokens per prompt. - generator (
torch.Generatororlist[torch.Generator], optional) — Seeds the sampled expansion; a list must matchprompt’s length, one generator per item.generatedraws from the global RNG, so the global RNG is seeded from this generator’s seed; laterrandn_tensorcalls keep usinggeneratordirectly.
Returns
str or list[str]
The expanded prompt(s).
Rewrites short prompts into detailed video+audio prompts with the Qwen2.5-VL text encoder, grounding them on
the reference image when one is given. A staticmethod so it can be used standalone, before running the
pipeline.
prepare_audio_latents
< source >( batch_size: intnum_channels_latents: intaudio_length: intdtype: dtypedevice: devicegenerator: typing.Union[torch.Generator, list[torch.Generator], NoneType]audio_latents: typing.Optional[torch.Tensor] = None )
Returns audio latents in the transformer’s (batch_size, audio_length, channels) layout. A user-provided audio_latents tensor is expected in the (batch_size, channels, audio_length) layout.
prepare_latents
< source >( batch_size: intnum_channels_latents: intheight: intwidth: intnum_frames: intdtype: dtypedevice: devicegenerator: typing.Union[torch.Generator, list[torch.Generator], NoneType]latents: typing.Optional[torch.Tensor] = None )
Returns video latents in the transformer’s (batch_size, num_frames, height, width, channels) layout.
A user-provided latents tensor is expected in the (batch_size, channels, num_frames, height, width) layout.
Kandinsky6SRPipeline
class diffusers.Kandinsky6SRPipeline
< source >( transformer: Kandinsky6SRTransformer3DModelvae: Kandinsky6SRVAEscheduler: diffusers.schedulers.scheduling_flow_match_euler_discrete.FlowMatchEulerDiscreteScheduler | diffusers.schedulers.scheduling_piflow.PiflowSchedulerlatent_upscaler: diffusers.models.latent_upscaler.latent_upscaler_kandinsky6_sr.Kandinsky6SRLatentUpscalerBank | None = None )
Parameters
- transformer (Kandinsky6SRTransformer3DModel) — Transformer that refines the latent tiles.
- vae (Kandinsky6SRVAE) — Causal video K-VAE used to encode the input video and decode the refined tiles.
- scheduler (FlowMatchEulerDiscreteScheduler or PiflowScheduler) —
Scheduler used with
transformerto denoise the tiles. Distilled checkpoints ship with a PiflowScheduler. - latent_upscaler (Kandinsky6SRLatentUpscalerBank, optional) — Latent upscalers for the supported scales.
Pipeline for video super-resolution with Kandinsky 6.
The video is split into overlapping spatial tiles, every tile is refined by the SR transformer at one of the tile
sizes the model was trained on (transformer.config.tile_sizes), and the refined tiles are blended back with Hann
windows. When the pipeline has a latent_upscaler, the tiles are cut from the K-VAE latents of the whole video and
upscaled in latent space; otherwise the pixel tiles are bilinearly upscaled and encoded.
This model inherits from DiffusionPipeline. Check the superclass documentation for the generic methods implemented for all pipelines (downloading, saving, running on a particular device, etc.).
__call__
< source >( video: typing.Union[list[PIL.Image.Image], list[list[PIL.Image.Image]], numpy.ndarray, torch.Tensor]resolution_scale: float = 2.25num_inference_steps: int = 4timesteps: list[int] | None = Nonesigmas: list[float] | None = Nonelq_noise_scale: float = 0.7min_overlap: float = 0.2tiles_batch_size: int = 1generator: typing.Union[torch.Generator, list[torch.Generator], NoneType] = Noneoutput_type: str = 'pil'return_dict: bool = True ) → Kandinsky6SRPipelineOutput or tuple
Parameters
- video (
list[PIL.Image.Image],np.ndarrayortorch.Tensor) — The low-resolution video(s), in any format preprocess_video() accepts, with1 + k * 4frames. Sizes are rounded down to a multiple of the VAE spatial factor. - resolution_scale (
float, defaults to2.25) — Total spatial upscale:2,4, or2.25(a 1.125x bilinear pre-upscale followed by the 2x path). - num_inference_steps (
int, defaults to4) — The number of denoising steps per tile. Use2with the distilled checkpoints. - timesteps (
list[int], optional) — Custom timesteps for schedulers that support them. - sigmas (
list[float], optional) — Custom sigmas for schedulers that support them. - lq_noise_scale (
float, defaults to0.7) — Amount of Gaussian noise mixed into the low-resolution latents (variance preserving) before denoising. - min_overlap (
float, defaults to0.2) — Minimum overlap between neighbouring tiles as a fraction of the tile size. - tiles_batch_size (
int, defaults to1) — Number of tiles denoised per transformer call. - generator (
torch.Generatororlist[torch.Generator], optional) — Generator(s) used for the noise mixed into the tiles. - output_type (
str, defaults to"pil") — The output format of the generated video:"pil","np"or"pt". - return_dict (
bool, defaults toTrue) — Whether or not to return a Kandinsky6SRPipelineOutput instead of a plain tuple.
Returns
Kandinsky6SRPipelineOutput or tuple
The super-resolved video; a one-element tuple when return_dict=False.
The call function to the pipeline for super-resolution.
Examples:
>>> import torch
>>> from diffusers import Kandinsky6SRPipeline, Kandinsky6TI2VAPipeline
>>> from diffusers.utils import export_to_video
>>> pipe = Kandinsky6TI2VAPipeline.from_pretrained(
... "kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", torch_dtype=torch.bfloat16
... )
>>> pipe.enable_model_cpu_offload()
>>> video = pipe(
... prompt="A cat and a dog baking a cake together in a kitchen.",
... height=480,
... width=864,
... num_inference_steps=16,
... guidance_scale=1.0,
... sample_audio=False,
... ).frames[0]
>>> sr_pipe = Kandinsky6SRPipeline.from_pretrained(
... "kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", torch_dtype=torch.bfloat16
... )
>>> # The transformer always runs attention through the `flex` backend; compiling avoids the eager
>>> # fallback's much higher memory use at video resolutions.
>>> sr_pipe.transformer.compile_repeated_blocks(fullgraph=True)
>>> sr_pipe.enable_model_cpu_offload()
>>> output = sr_pipe(video=video, resolution_scale=2.25, num_inference_steps=2)
>>> export_to_video(output.frames[0], "output_sr.mp4", fps=24)Decodes scaled K-VAE latents into a (batch_size, channels, num_frames, height, width) video in [-1, 1].
Encodes a (batch_size, channels, num_frames, height, width) video in [-1, 1] into K-VAE latents scaled by
the VAE scaling_factor.
Kandinsky6SRLatentUpscalerBank
class diffusers.Kandinsky6SRLatentUpscalerBank
< source >( in_channels: int = 64stage_channels: tuple[int, int, int] = (2048, 1024, 512)num_pre_blocks: int = 5num_mid_blocks: int = 3num_post_blocks: int = 3num_x2_adapter_blocks: int = 2scales: tuple[int, ...] = (2, 4)scaling_factor: float = 0.910344004631042 )
Parameters
- in_channels (
int, defaults to64) — Number of latent channels. - stage_channels (
tuple[int, int, int], defaults to(2048, 1024, 512)) — Feature widths of the three stages of the cascade. - num_pre_blocks (
int, defaults to5) — Residual blocks before the first upsample of the x4 model. - num_mid_blocks (
int, defaults to3) — Residual blocks between the two upsamples. - num_post_blocks (
int, defaults to3) — Residual blocks after the last upsample. - num_x2_adapter_blocks (
int, defaults to2) — Residual blocks of the x2 model’s adapter. - scales (
tuple[int, ...], defaults to(2, 4)) — Spatial scales the bank provides an upscaler for. - scaling_factor (
float, defaults to0.910344) — Scale the input latents are expected to carry (the K-VAEscaling_factor).
Bank of latent upscalers used by Kandinsky6SRPipeline: one Kandinsky6SRLatentUpscaler per supported spatial
scale, operating on K-VAE latents.
forward
< source >( latents: Tensorscale: intreturn_dict: bool = True )
Parameters
- latents (
torch.Tensorof shape(batch_size, in_channels, num_frames, height, width)) — K-VAE latents scaled byscaling_factor. - scale (
int) — Spatial upscale factor; one ofscales. - return_dict (
bool, defaults toTrue) — Whether to return a~models.autoencoder_kl.DecoderOutputinstead of a plain tuple.
Kandinsky6TI2VAPipelineOutput
class diffusers.Kandinsky6TI2VAPipelineOutput
< source >( frames: typing.Union[torch.Tensor, numpy.ndarray, list[list[PIL.Image.Image]]]audio: typing.Union[torch.Tensor, numpy.ndarray, NoneType] = None )
Parameters
- frames (
torch.Tensor,np.ndarray, orlist[list[PIL.Image.Image]]) — The generated video. A nested list of lengthbatch_sizeholdingnum_framesPIL images each, or a NumPy array or torch tensor of shape(batch_size, num_frames, height, width, channels)/(batch_size, num_frames, channels, height, width). Withoutput_type="latent", the video latents of shape(batch_size, channels, num_latent_frames, latent_height, latent_width). - audio (
torch.Tensorornp.ndarray, optional) — The generated waveforms of shape(batch_size, num_samples)in[-1, 1]at the audio VAE’s sample rate, orNonewhen audio was not sampled. Withoutput_type="latent", the audio latents of shape(batch_size, channels, audio_length).
Output class for Kandinsky6TI2VAPipeline.
Kandinsky6SRPipelineOutput
class diffusers.Kandinsky6SRPipelineOutput
< source >( frames: typing.Union[torch.Tensor, numpy.ndarray, list[list[PIL.Image.Image]]] )
Parameters
- frames (
torch.Tensor,np.ndarray, orlist[list[PIL.Image.Image]]) — The super-resolved video. A nested list of lengthbatch_sizeholdingnum_framesPIL images each, or a NumPy array or torch tensor of shape(batch_size, num_frames, height, width, channels)/(batch_size, num_frames, channels, height, width).
Output class for Kandinsky6SRPipeline.