Fix or replace part of an existing video in one pass: SAM 3.1 on ComfyUI's core nodes tracks the thing you name, the mask is grown to cover its edge, and VACE regenerates inside it from a prompt. The mask never leaves the graph, so there is no intermediate clip to write, upload and read back.
Tags: motionsam3sam3.1vaceinpaintcompositionrunpod-serverlessvolume-backed
Inputs (19)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
source_videovideorequiredtrack_promptstringrequireddefault personpromptstringrequireddefault the same scene with the tracked subject replaced, cinematic lightingnegative_promptstringdefault 色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走reference_imageimagemask_growintegerdefault 12min -64max 64track_thresholdfloatdefault 0.5min 0.0max 1.0track_max_objectsintegerdefault 4min 0max 64track_detect_intervalintegerdefault 1min 1max 240strengthfloatdefault 1.0min 0.0max 2.0resolutionenumdefault 480p_landscape_832x480480p_landscape_832x480480p_portrait_480x832720p_landscape_1280x720720p_portrait_720x1280lengthintegerdefault 81min 5max 241fpsfloatdefault 16.0min 1.0max 30.0stepsintegerdefault 20min 1max 60cfgfloatdefault 5.0min 1.0max 20.0shiftfloatdefault 8.0min 0.0max 20.0seedintegerdefault 0min 0samplerenumdefault uni_pcuni_pceulerdpmpp_2mschedulerenumdefault simplesimplenormalbetakarrasComfyUI node graph (19)#
The executable ComfyUI prompt graph: 19 nodes across 17 distinct node classes, wired by 25 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (19)#
1LoadVideocorefile = {{source_video}} tmplVIDEO2GetVideoComponentscorevideo = ◂ node 1 · out[0]IMAGEAUDIOFLOATCOMBOCOMBO3CheckpointLoaderSimplecoreckpt_name = sam3.1_multiplex_fp16.safetensorsMODELCLIPVAE4CLIPTextEncodecoreclip = ◂ node 3 · out[1]text = {{track_prompt}} tmplCONDITIONING5SAM3_VideoTrackcoreimages = ◂ node 2 · out[0]model = ◂ node 3 · out[0]conditioning = ◂ node 4 · out[0]detection_threshold = {{track_threshold}} tmplmax_objects = {{track_max_objects}} tmpldetect_interval = {{track_detect_interval}} tmplSAM3_TRACK_DATA6SAM3_TrackToMaskcoretrack_data = ◂ node 5 · out[0]object_indices = MASK7GrowMaskcoremask = ◂ node 6 · out[0]expand = {{mask_grow}} tmpltapered_corners = trueMASK8UNETLoadercoreunet_name = wan2.1_vace_14B_fp16.safetensorsweight_dtype = defaultMODEL9CLIPLoadercoreclip_name = umt5_xxl_fp16.safetensorstype = wandevice = defaultCLIP10VAELoadercorevae_name = wan_2.1_vae.safetensorsVAE11ModelSamplingSD3coremodel = ◂ node 8 · out[0]shift = {{shift}} tmplMODEL12CLIPTextEncodecoreclip = ◂ node 9 · out[0]text = {{prompt}} tmplCONDITIONING13CLIPTextEncodecoreclip = ◂ node 9 · out[0]text = {{negative_prompt}} tmplCONDITIONING14WanVaceToVideocorepositive = ◂ node 12 · out[0]negative = ◂ node 13 · out[0]vae = ◂ node 10 · out[0]width = {{resolution_map[resolution].width}} tmplheight = {{resolution_map[resolution].height}} tmpllength = {{length}} tmplbatch_size = 1strength = {{strength}} tmplcontrol_video = ◂ node 2 · out[0]control_masks = ◂ node 7 · out[0]CONDITIONINGCONDITIONINGLATENTINT15KSamplercoremodel = ◂ node 11 · out[0]positive = ◂ node 14 · out[0]negative = ◂ node 14 · out[1]latent_image = ◂ node 14 · out[2]seed = {{seed}} tmplsteps = {{steps}} tmplcfg = {{cfg}} tmplsampler_name = {{sampler}} tmplscheduler = {{scheduler}} tmpldenoise = 1.0LATENT16TrimVideoLatentcoresamples = ◂ node 15 · out[0]trim_amount = ◂ node 14 · out[3]LATENT17VAEDecodecoresamples = ◂ node 16 · out[0]vae = ◂ node 10 · out[0]IMAGE18CreateVideocoreimages = ◂ node 17 · out[0]fps = {{fps}} tmplVIDEO19SaveVideocorevideo = ◂ node 18 · out[0]filename_prefix = sam3-vaceformat = mp4Parameter banks (3)#
The prompt / configuration lookup tables this workflow keys into from its inputs — the vocabulary that turns a style / palette / preset selection into graph parameters.
resolution_map (4)#
480p_landscape_832x480{"width": 832, "height": 480}480p_portrait_480x832{"width": 480, "height": 832}720p_landscape_1280x720{"width": 1280, "height": 720}720p_portrait_720x1280{"width": 720, "height": 1280}post_render (1)#
optional_image_input{"input": "reference_image", "target_class": "WanVaceToVideo", "target_input": "reference_image"}requires_families (3)#
sam31wan21-vacewan-sharedModels & dependencies#
Models required (4)#
sam3.1_multiplex_fp16.safetensorswan2.1_vace_14B_fp16.safetensorsumt5_xxl_fp16.safetensorswan_2.1_vae.safetensorsOutput contract#
What a successful run of this workflow returns.
primary{"type": "video", "format": "mp4", "codec": "h264", "fps_source": "declared", "audio": false, "alpha": false, "description": "The clip with the tracked region regenerated."}Taxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyregion-edited-videooutputPackageProfilevideo-master-profilecontrolModalitiesmodel-lockprompt-template-locksampler-scheduler-lockseed-locktemporal-lockcontrolnet-segmentationcontrolnet-inpaintreference-ensembleconsistencyDimensionsidentitymotionlightingenvironmentnotesThe composition of sam3-video-track and wan21-vace-edit's inpaint mode in ONE graph: the MASK goes straight from SAM3_TrackToMask through GrowMask into WanVaceToVideo, so there is no intermediate mask clip to write, upload and read back.