视频换人
视频换人组合 GPU API 的请求、阶段、轮询与错误说明
文档更新时间:2026年09月08日 12时53分18秒(UTC+8)
视频换人(组合 API)
最后更新:2026-09-07
| 项目 | 值 |
|---|---|
model | video-person-swap |
| 输入 | 1 个源视频 + 1–2 张人物图 |
| 内部 Scheduler | POST /v1/video-person-swap |
| 输出 | 1 张可提前读取的中间 PNG + 1 个经过媒体校验的视频 |
组合流程
阶段 1:首帧提取 → BFS 换头 → BodySwap 换身体 → 持久化新首帧
阶段 2:将双阶段预编辑的新首帧送入 Wan2.1 SCAIL2
视频换人首帧固定复用图片工具箱 krea2-person-swap-v1 的身份替换约束。BFS 与 BodySwap 两个预编辑阶段都沿用源首帧的竖版尺寸链,不再先生成 1:1 方图,也不再在 BodySwap 前强制缩放为 1024×1024。当前竖版工作流的中间结果为 512×896;BFS 使用 8 步、ref_boost=1、grounding_px=512 的身份保持基线。接口返回的“中间结果”是 BFS 换头并继续完成 BodySwap 后、送入 SCAIL2 前的新首帧,不是仅完成 BFS 的临时图。
当前生产实现把 BFS、BodySwap 和 SCAIL2 编排在同一个组合工作图中,对调用方只暴露一个任务 ID。这里的“双阶段图片换人”是 BFS 与 BodySwap 两个连续预编辑阶段,不是两个独立 API 任务。中间 PNG 在 BodySwap 输出节点生成;Scheduler 查询一旦返回 intermediate_image_available: true,调用方即可下载,不必等最终视频完成。
只有一张人物图时,换头和换身体阶段复用该图。第二张图存在时,第 1 张用于头部,第 2 张用于身体。
默认提示词
提示词留空时,视频换脸和视频换人使用以下人物替换约束;它只替换人物并要求保留源视频场景,不把“人物在跳舞”当作动作或背景重绘指令:
Replace only the person in the source video with the person from the reference image. Preserve the reference person's facial identity, hairstyle, body shape, proportions, and clothing. Preserve the source video's original motion, pose sequence, facial motion, background, environment, objects, lighting, camera movement, framing, composition, timing, and duration. Do not modify or regenerate the scene. Keep the replacement person's identity, body shape, and clothing consistent across all frames.用户提交非空提示词时仍原样使用用户内容。video-undress 与 scail2-person-swap 的空提示默认值保持 the woman is dancing。
创建任务
curl -X POST "$BASE_URL/api/proxy/v1/videos" \
-H "Cookie: ic_session=$IC_SESSION" \
-H "X-IC-Operation-ID: $(uuidgen)" \
-F "model=video-person-swap" \
-F "prompt=Replace only the person in the source video with the person from the reference image. Preserve the reference person's facial identity, hairstyle, body shape, proportions, and clothing. Preserve the source video's original motion, pose sequence, facial motion, background, environment, objects, lighting, camera movement, framing, composition, timing, and duration. Do not modify or regenerate the scene. Keep the replacement person's identity, body shape, and clothing consistent across all frames." \
-F "seed=12345" \
-F "input_reference[]=@face.png" \
-F "input_reference[]=@body.png" \
-F "input_reference[]=@source.mp4"源视频支持 MP4/MOV、最大 500MB,必须声明 video/mp4 或 video/quicktime 并含可解码音轨;无声素材应先封装静音 AAC 音轨。人物图支持 JPG/PNG/BMP/WEBP、单张最大 30MB,必须声明真实 image/* MIME。Proxy 会按 MIME 分类,不把文件名后缀当作唯一判断,也不接受 application/octet-stream 代替真实媒体类型。
状态与结果
- 创建成功保存返回的
id。 POST /api/proxy/v1/videos/{id}轮询pending/running/completed/failed;运行中即可通过响应里的intermediate_image_available判断新首帧是否已生成。- 完成后调用
GET /api/proxy/v1/videos/{id}/content下载视频。 - 取消调用
POST /api/proxy/v1/videos/{id}/cancel;只有服务端确认取消后才视为终止。
完整 Python 调用
下面代码可直接保存为 call_video_person_swap.py。先执行 pip install requests,再设置 BASE_URL、IC_SESSION 并准备示例媒体。代码会提交任务、轮询、下载最终 MP4;若存在首帧预处理结果,也会下载为 intermediate.png。
import mimetypes
import os
import time
import uuid
from contextlib import ExitStack
from pathlib import Path
import requests
BASE_URL = os.environ.get("BASE_URL", "https://ic.xshow.live").rstrip("/")
IC_SESSION = os.environ["IC_SESSION"]
MODEL = "video-person-swap"
POLL_SECONDS = 5
TIMEOUT_SECONDS = 6 * 60 * 60
DEFAULT_PROMPT = "Replace only the person in the source video with the person from the reference image. Preserve the reference person's facial identity, hairstyle, body shape, proportions, and clothing. Preserve the source video's original motion, pose sequence, facial motion, background, environment, objects, lighting, camera movement, framing, composition, timing, and duration. Do not modify or regenerate the scene. Keep the replacement person's identity, body shape, and clothing consistent across all frames."
session = requests.Session()
session.cookies.set("ic_session", IC_SESSION)
def mime(path: str) -> str:
value = mimetypes.guess_type(path)[0]
if not value:
raise ValueError(f"无法识别媒体 MIME: {path}")
return value
def checked(response: requests.Response) -> requests.Response:
if not response.ok:
raise RuntimeError(f"HTTP {response.status_code}: {response.text[:1000]}")
return response
with ExitStack() as stack:
files = [
("input_reference[]", (Path("face.png").name, stack.enter_context(open("face.png", "rb")), mime("face.png"))),
("input_reference[]", (Path("body.png").name, stack.enter_context(open("body.png", "rb")), mime("body.png"))),
("input_reference[]", (Path("source.mp4").name, stack.enter_context(open("source.mp4", "rb")), mime("source.mp4")))
]
created = checked(session.post(
f"{BASE_URL}/api/proxy/v1/videos",
headers={"X-IC-Operation-ID": str(uuid.uuid4())},
data={
"model": MODEL,
"prompt": DEFAULT_PROMPT,
"seed": "12345",
},
files=files,
timeout=1800,
)).json()
task_id = created.get("id") or (created.get("data") or {}).get("id")
if not task_id:
raise RuntimeError(f"创建响应缺少任务 ID: {created}")
print("task_id:", task_id)
deadline = time.time() + TIMEOUT_SECONDS
while time.time() < deadline:
state = checked(session.post(
f"{BASE_URL}/api/proxy/v1/videos/{task_id}",
timeout=60,
)).json()
status = str(state.get("status") or (state.get("data") or {}).get("status") or "pending").lower()
print("status:", status)
if status == "completed":
video = checked(session.get(
f"{BASE_URL}/api/proxy/v1/videos/{task_id}/content",
timeout=1800,
))
Path("result.mp4").write_bytes(video.content)
print("saved: result.mp4")
payload = state.get("data") if isinstance(state.get("data"), dict) else state
if payload.get("intermediate_image_available") is True:
intermediate = checked(session.get(
f"{BASE_URL}/api/proxy/v1/videos/{task_id}/intermediate",
timeout=600,
))
Path("intermediate.png").write_bytes(intermediate.content)
print("saved: intermediate.png")
break
if status in {"failed", "cancelled", "canceled", "expired"}:
raise RuntimeError(f"任务失败: {state}")
time.sleep(POLL_SECONDS)
else:
raise TimeoutError(f"任务 {task_id} 超过 {TIMEOUT_SECONDS} 秒仍未完成")阶段失败
错误会保留首帧提取、图片预编辑、SCAIL2 提交、GPU 执行或结果校验阶段的真实原因。400 通常是媒体数量/格式错误,503 是工作流未就绪,502 是上游或结果媒体校验失败。