最典型的爆显存:KSampler 阶段 torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate ... / 全站资源搜索
搜索 更新 2026-10-02
报错 012 显存与内存不足(OOM) 有来源可查

最典型的爆显存:KSampler 阶段 torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate ...

报错原文(原样)
!!! Exception during processing !!! CUDA error: out of memory
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
Traceback (most recent call last):
File "C:\ComfyUI-Zluda\execution.py", line 524, in execute
output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
File "C:\ComfyUI-Zluda\comfy\samplers.py", line 967, in inner_sample
if latent_image is not None and torch.count_nonzero(latent_image) > 0: #Don't shift the empty latent image.
RuntimeError: CUDA error: out of memory
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
检索关键词 CUDA out of memory、CUDA error: out of memory、Tried to allocate、torch.cuda.OutOfMemoryError、CUDA kernel errors might be asynchronously reported、Exception during processing
先看这一句 排队出图后控制台刷红字,进度条卡在 KSampler(或 VAE Decode)阶段不动,ComfyUI 面板提示 Prompt execution failed,图像完全不生成。
适用版本 全版本通用;报错发生在采样阶段的激活值分配,与 ComfyUI 版本无强绑定。 · 核实日期 2026-09-26 · 状态 已核实
适用环境:Windows 便携版、秋叶/绘世整合包、ComfyUI Desktop、Linux + 云 GPU、Docker/服务器

现象与触发场景

常见触发场景:

  • 8GB 显卡跑 SDXL / Flux 的默认工作流
  • 把分辨率从 1024x1024 提到 1536 以上,或把 batch size 从 1 提到 2 以上
  • 同一台机器同时开着游戏、浏览器硬件加速或 OBS 等占用显存的程序
  • 云 GPU 实例上已被同实例的其他进程占用了一部分显存

根因

  • ComfyUI 需要在显存里同时容纳 UNet/DiT 权重、文本编码器、VAE,以及采样过程中的激活值(attention 中间张量),激活值随分辨率与 batch 呈平方级增长,超出的那一刻由 PyTorch CUDA 分配器抛出 OOM。
  • ComfyUI 默认会为系统/其他软件保留一部分显存(EXTRA_RESERVED_VRAM),Windows 上因为共享显存机制默认保留得更多(600MB 而非 400MB),所以可用显存比 nvidia-smi 显示的数字要小。
  • 上层包装(如 ZLUDA、ROCm、DirectML 后端)会把同一错误以 `RuntimeError: CUDA error: out of memory` 或 `HIP out of memory` 形式抛出,内核报错还可能被异步延迟到后续 API 调用时才显形。

解决步骤

已按「最可能有效」排序。请从上往下做,每做完一步用「验证」确认,不要一次改五处。

  1. 先降到能跑通的输入规模,确认是容量而不是别的问题
    把分辨率减半(如 1024 -> 768 或 512)、batch size 改回 1、总帧数/时长减半,先跑通一次。能跑通说明纯粹是显存容量不够,再按下面的步骤逐项加回来。官方排障文档明确把“降低分辨率/batch size”列为 OOM 的第一优先动作。
    # 图像:Empty Latent Image 的 width/height 从 1024x1024 改为 768x768;Batch Size 改为 1
    # 视频:帧数(length/frames)先减半,如 81 -> 33
  2. 给系统和其他程序预留固定显存(Windows 上尤其有效)
    在启动命令追加 --reserve-vram,单位是 GB。ComfyUI 会把这部分显存从可用池里扣除,避免算到一半被系统或浏览器抢走触发 OOM。官方排障文档给出的示例就是 2GB。
    python main.py --reserve-vram 2
    # Windows 便携版:编辑 run_nvidia_gpu.bat,在 .\python_embeded\python.exe -s .\ComfyUI\main.py 后面追加参数
  3. 改用更省显存的注意力实现并关闭智能缓存
    --use-pytorch-cross-attention 使用 PyTorch 2.0 的原生注意力(源码 help:Use the new pytorch 2.0 cross attention function);--use-split-cross-attention 用分块注意力换显存;--disable-smart-memory 强制 ComfyUI 主动把模型卸载回普通内存,而不是尽量留在显存里。AMD 用户的官方 issue #11624 实测 --use-pytorch-cross-attention --disable-smart-memory 组合能显著减少崩溃。
    python main.py --use-pytorch-cross-attention --disable-smart-memory
    python main.py --use-split-cross-attention
    python main.py --reserve-vram 2 --use-pytorch-cross-attention --disable-smart-memory
  4. 把 VAE 解码单独降级处理(解码阶段是爆显存第二大高发点)
    VAE Decode 阶段显存峰值往往远高于采样阶段。可换成内置的 VAE Decode (Tiled) 节点分块解码,或启用 --fp16-vae / --bf16-vae 降低 VAE 精度;实在不行用 --cpu-vae 把 VAE 放 CPU(很慢但一定能过)。
    # 节点替换:VAEDecode -> VAEDecodeTiled(内置节点,可调 tile_size)
    --fp16-vae
    --bf16-vae
    --cpu-vae

验证是否修好

重跑同一工作流,控制台不再出现 CUDA out of memory,图片正常保存到 output 目录;再用 nvidia-smi 观察峰值显存占用,确认留有 500MB 以上余量。

补充说明

源码核实:comfy/cli_args.py 中 --reserve-vram 的 help 为 'Set the amount of vram in GB you want to reserve for use by your OS/other software. By default some amount is reserved depending on your OS.';comfy/model_management.py 中 EXTRA_RESERVED_VRAM 在 Windows 下为 600MB,注释原文为 '#Windows is higher because of the shared vram issue',Linux 为 400MB。

来源

按可信度排列:官方 issue / 官方文档 > 节点仓库 issue > 社区帖 > 中文社区文章。 链接以纯文本给出(本站不做站外跳转),需要核对时请自行复制到浏览器打开。
  1. [34] (官方 issue) pls help ive been struggling for days error after error getting ran ou https://github.com/Comfy-Org/ComfyUI/issues/12713
  2. [6] (官方文档) How to Troubleshoot and Solve ComfyUI Issues - ComfyUI Documentation https://docs.comfy.org/troubleshooting/overview
  3. [35] (官方文档) ComfyUI/comfy/cli_args.py(参数与 help 原文) https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/cli_args.py
  4. [36] (官方文档) ComfyUI/comfy/model_management.py(EXTRA_RESERVED_VRAM 等) https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/model_management.py

这条报错要用到的东西

按本条报错所属的类别(显存与内存不足(OOM))列出站内相关的下载包与板块入口。 它们不是「必须下载才能修好」,而是这一类问题里最常要动到的东西;具体怎么做,以上面的解决步骤为准。

ComfyUI 权重库(167 个文件 / 约 1187 GB)按 ComfyUI 的 models 目录分好类的 167 个权重文件 / 每个文件一份说明(放哪、被哪些工作流引用) 待挂载
权重库同一个模型常有更小的量化版本,换一个能省一大截显存 板块 工作流换一条依赖更轻的工作流 板块

同一类别的其他条目

这个解法对你有效吗 记录只存在你自己的浏览器里,不需要注册账号。

返回报错库