Wan2.2-Animate-2: a reference still performs the motion of a driving clip, end to end. The driving frames go to the model directly (no ONNX pose/face preprocess, unlike wan22-animate); the character comes from the reference and the background and camera from the prompt. One 81-frame window of the ComfyUI Animate-2 distilled template.
Tags: motionwan2.2animate-2performance-transferrunpod-serverlessvolume-backed
Inputs (13)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
reference_imageimagerequireddriving_videovideorequiredpromptstringrequireddefault Character appearance description: the character in the reference image.
Background description: Plain light gray studio background, soft even lighting, no decorations.motion_promptstringdefault a person dancingnegative_promptstringdefault 色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走speed_modeenumdefault distilleddistilledresolutionenumdefault 482x854482x854854x482480x832832x480512x896896x512720x12801280x720lengthintegerdefault 81min 5max 81pose_strengthfloatdefault 1.0min 0.0max 10.0pose_start_percentfloatdefault 0.0min 0.0max 1.0pose_end_percentfloatdefault 1.0min 0.0max 1.0reference_image_strengthfloatdefault 1.0min 0.0max 10.0seedintegerdefault 0min 0ComfyUI node graph (25)#
The executable ComfyUI prompt graph: 25 nodes across 21 distinct node classes, wired by 36 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (25)#
1UNETLoadercoreunet_name = {{speed_unet_map[speed_mode]}} tmplweight_dtype = defaultMODEL2CLIPLoadercoreclip_name = umt5_xxl_fp16.safetensorstype = wandevice = defaultCLIP3VAELoadercorevae_name = wan_2.1_vae.safetensorsVAE4CLIPVisionLoadercoreclip_name = clip_vision_h.safetensorsCLIP_VISION5WanAnimate2Cachecoremodel = ◂ node 1 · out[0]device = gpudtype = int8MODEL6ModelSamplingSD3coremodel = ◂ node 5 · out[0]shift = 5MODEL7CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{prompt}} tmplCONDITIONING8CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{negative_prompt}} tmplCONDITIONING9CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{motion_prompt}} tmplCONDITIONING10LoadImagecoreimage = {{reference_image}} tmplIMAGEMASK11LoadVideocorefile = {{driving_video}} tmplVIDEO12GetVideoComponentscorevideo = ◂ node 11 · out[0]IMAGEAUDIOFLOATCOMBOCOMBO13ResizeImageMaskNodecoreinput = ◂ node 12 · out[0]resize_type = scale dimensionsresize_type.width = {{resolution_map[resolution].width}} tmplresize_type.height = {{resolution_map[resolution].height}} tmplresize_type.crop = centerscale_method = areaIMAGE14ResizeImageMaskNodecoreinput = ◂ node 10 · out[0]resize_type = scale dimensionsresize_type.width = {{resolution_map[resolution].width}} tmplresize_type.height = {{resolution_map[resolution].height}} tmplresize_type.crop = centerscale_method = areaIMAGE15ImageFromBatchcoreimage = ◂ node 13 · out[0]batch_index = 0length = 1IMAGE16CLIPVisionEncodecoreclip_vision = ◂ node 4 · out[0]image = ◂ node 14 · out[0]crop = noneCLIP_VISION_OUTPUT17CLIPVisionEncodecoreclip_vision = ◂ node 4 · out[0]image = ◂ node 15 · out[0]crop = noneCLIP_VISION_OUTPUT18WanAnimate2ToVideocorepositive = ◂ node 7 · out[0]negative = ◂ node 8 · out[0]vae = ◂ node 3 · out[0]reference_image = ◂ node 14 · out[0]pose_video = ◂ node 13 · out[0]clip_vision_output = ◂ node 16 · out[0]positive_pose = ◂ node 9 · out[0]clip_vision_output_pose = ◂ node 17 · out[0]width = {{resolution_map[resolution].width}} tmplheight = {{resolution_map[resolution].height}} tmpllength = {{length}} tmplbatch_size = 1video_frame_offset = 0pose_strength = {{pose_strength}} tmplpose_start_percent = {{pose_start_percent}} tmplpose_end_percent = {{pose_end_percent}} tmplreference_image_strength = {{reference_image_strength}} tmplCONDITIONINGCONDITIONINGLATENTINTINTINT19KSamplerSelectcoresampler_name = {{speed_sampler_map[speed_mode]}} tmplSAMPLER20BasicSchedulercoremodel = ◂ node 5 · out[0]scheduler = simplesteps = {{speed_steps_map[speed_mode]}} tmpldenoise = 1.0SIGMAS21SamplerCustomcoremodel = ◂ node 6 · out[0]add_noise = truenoise_seed = {{seed}} tmplcfg = {{speed_cfg_map[speed_mode]}} tmplpositive = ◂ node 18 · out[0]negative = ◂ node 18 · out[1]sampler = ◂ node 19 · out[0]sigmas = ◂ node 20 · out[0]latent_image = ◂ node 18 · out[2]LATENTLATENT22TrimVideoLatentcoresamples = ◂ node 21 · out[0]trim_amount = ◂ node 18 · out[3]LATENT23VAEDecodecoresamples = ◂ node 22 · out[0]vae = ◂ node 3 · out[0]IMAGE24CreateVideocoreimages = ◂ node 23 · out[0]audio = ◂ node 12 · out[1]fps = ◂ node 12 · out[2]VIDEO25SaveVideocorevideo = ◂ node 24 · out[0]filename_prefix = wan22-animate-2format = mp4Parameter banks (6)#
The prompt / configuration lookup tables this workflow keys into from its inputs — the vocabulary that turns a style / palette / preset selection into graph parameters.
resolution_map (8)#
482x854{"width": 482, "height": 854}854x482{"width": 854, "height": 482}480x832{"width": 480, "height": 832}832x480{"width": 832, "height": 480}512x896{"width": 512, "height": 896}896x512{"width": 896, "height": 512}720x1280{"width": 720, "height": 1280}1280x720{"width": 1280, "height": 720}speed_unet_map (1)#
distilledwan_animate_2_distill_bf16.safetensorsspeed_steps_map (1)#
distilled10speed_sampler_map (1)#
distilledlcmspeed_cfg_map (1)#
distilled1.0requires_families (1)#
wan22-animate2Models & dependencies#
Models required (4)#
wan_animate_2_distill_bf16.safetensorsumt5_xxl_fp16.safetensorswan_2.1_vae.safetensorsclip_vision_h.safetensorsOutput contract#
What a successful run of this workflow returns.
primary{"type": "video", "format": "mp4", "codec": "h264", "fps_source": "source", "audio": true, "alpha": false, "description": "The reference character performing the driving motion, at the driving clip's frame rate with its audio track when it has one."}Taxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyanim-character-turnaround-packoutputPackageProfilevideo-master-profilecontrolModalitiesmodel-lockprompt-template-locksampler-scheduler-lockseed-lockidentity-lockpose-constrainttemporal-lockconsistencyDimensionsidentitymotionenvironmentwardrobenotesWan2.2-Animate-2 (A.01.06). pose-constraint, not openpose-sequence: the driving frames themselves are the pose video, with no skeleton drawn. identity-lock is the reference latent plus its CLIP-vision encoding; environment comes from the prompt's background description.