来源:D3D12RHI 源码阅读笔记
引擎版本:UE 5.8
面向读者:UE 图形程序员
UE 在 Windows 上创建的 D3D12 heap 和 committed resource,默认带着 D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT 出生:刚创建完,GPU 还碰不了它。负责在 GPU 用到它之前把它换进来的,是每条 FD3D12CommandList 身上一个几乎没有存在感的成员 —— ResidencySet。
绝大多数资源经由 barrier 和 descriptor cache 自动登记进去,所以平时感觉不到它。可一旦某条录制路径漏了登记,出事的往往不是当帧,而是显存吃紧、LRU 开始驱逐之后的某一帧,现场是一个很难往回追的 DEVICE_REMOVED。
目录
- ResidencySet 是什么
- 为什么它是正确性前提
- 跟踪谁:对象模型与粒度
- 一条 CL 的完整生命周期
- 提交那一刻:ExecuteSubset 与 ProcessPagingWork
- 移除到底发生在哪一层
- Reserved Resource 的延迟登记
- 开关与常量
- 踩坑与排查
- 总结
- 参考链接
一、ResidencySet 是什么
FD3D12CommandList::ResidencySet 是这条 command list 在 GPU 上执行时会访问到的全部 ID3D12Pageable 的去重清单。CPU 录制时往里填,提交时交给 FD3D12ResidencyManager 消费:Manager 保证清单里的对象在 ExecuteCommandLists 之前全部 resident,同时把清单之外、长时间没被用到的对象换出去。
// Engine/Source/ThirdParty/Windows/D3DX12/Include/d3dx12residency.h:156
// This represents a set of objects which are referenced by a command list i.e. every time a resource
// is bound for rendering, clearing, copy etc. the set must be updated to ensure the it is resident
// for execution.
class ResidencySet| 性质 | 含义 |
|---|---|
| 按 CL 粒度 | 每个 FD3D12CommandList 对象一个,构造时创建,const 指针终生绑定 |
| 只增不删 | 录制期只有 Insert,没有针对单个对象的删除 |
| 一次性 | 每次 Open() 逻辑清零,提交之后内容即作废 |
| 不持有所有权 | 只存 ManagedObject* 裸指针数组 ppSet |
1.1 代码分布
底层是微软 D3D12 Residency Starter Library 的 d3dx12residency.h,Epic 的改动都用 BEGIN EPIC MOD / END EPIC MOD 标出;库的用法就是一条 CL 配一个 ResidencySet,提交时把两组数组一起交给 ResidencyManager::ExecuteCommandLists。UE 这一侧的薄封装在 D3D12Residency.h,每个包装函数都先判断 GEnableResidencyManagement,关掉后整套退化为空操作。
| 文件 | 目录 |
|---|---|
d3dx12residency.h(residency 库本体) | Engine/Source/ThirdParty/Windows/D3DX12/Include/ |
D3D12Residency.h、D3D12CommandList.*、D3D12Submission.*、D3D12Resources.*、D3D12Adapter.cpp、D3D12Device.*、D3D12CommandContext.*、各分配器 | Engine/Source/Runtime/D3D12RHI/Private/ |
WindowsD3D12Device.cpp | Engine/Source/Runtime/D3D12RHI/Private/Windows/ |
D3D12RHI.h、ID3D12DynamicRHI.h | Engine/Source/Runtime/D3D12RHI/Public/ |
二、为什么它是正确性前提
2.1 D3D12 把驻留管理交还给了应用
- D3D11 下,API 之下的层知道每个 command buffer 要用哪些资源,于是能替应用把要用的换入、把久未使用的换出。D3D12 的 bindless 和更直接的内存访问让下层拿不到这份信息,这件事只能由应用自己做。
- heap 和 committed resource 默认在创建流程的最后一步变成 resident。进程驻留量超出显存预算(
IDXGIAdapter3::QueryVideoMemoryInfo给出的 Budget)时,OS 会间歇性冻结进程,创建 API 也可能失败;内核把 heap 从显存挪到系统内存只是万不得已的兜底,不能指望。 - GPU 不支持缺页。未驻留的对象 GPU 既不能读也不能写,违反这条的典型现场是 GPU page fault 引发的
DEVICE_REMOVED。
2.2 UE:被跟踪的对象一律 EVICTED 出生
// Engine/Source/Runtime/D3D12RHI/Public/D3D12RHI.h:29(PLATFORM_WINDOWS 分支)
#define ENABLE_RESIDENCY_MANAGEMENT 1
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Adapter.cpp:38
bool GD3D12StartResourceResident = false;
static FAutoConsoleVariableRef CVarResourcesStartResident(
TEXT("D3D12.ResourcesStartResident"),
GD3D12StartResourceResident,
TEXT("When disabled (default): all D3D12 resources start EVICTED. GPU Residency will be managed during command list submission.")
TEXT("When enabled: all D3D12 resources are RESIDENT on the GPU at the time of creation."),
ECVF_ReadOnly
);“EVICTED 出生”由三处代码共同落实:
| 环节 | 位置 | 做法 |
|---|---|---|
| 创建 D3D 对象 | D3D12Resources.cpp:813(committed)、D3D12PoolAllocator.cpp:86 / :493、D3D12Allocation.cpp:303、D3D12TransientResourceAllocator.cpp:51 | heap flag 加上 D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT,跳过创建流程最后默认的 MakeResident |
| 初始化跟踪句柄 | D3D12Residency.h:63 | ManagedObject::Initialize(..., EVICTED),句柄状态直接记为 EVICTED |
| 开始跟踪 | D3D12Residency.h:111、d3dx12residency.h:875-920 | Manager 已 SetStartEvicted(true);BeginTrackingObject 把对象挂进 LRU 的 evicted 链表,若对象仍标记为 RESIDENT 则当场 Evict |
这一套依赖 OS 认识 CREATE_NOT_RESIDENT 这个 flag。UE 以能否拿到 ID3D12Device8 为准(D3D12Adapter.cpp:1101-1111),拿不到就退回创建即 resident:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Adapter.cpp:1138
#if ENABLE_RESIDENCY_MANAGEMENT
// If the OS does not understand D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT (pre-Windows 10 20H1) resources cannot be created evicted.
if (!bHeapCreateNotResidentSupported)
{
UE_LOGF(LogD3D12RHI, Log, "D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT is not supported (requires ID3D12Device8); forcing resources to start resident.");
GD3D12StartResourceResident = true;
}
#endif // ENABLE_RESIDENCY_MANAGEMENT✅ 由此得出最关键的一条:ResidencySet 不是性能优化,而是正确性前提。一个 EVICTED 出生的跟踪对象,只要从没进过任何一条 CL 的 ResidencySet,就永远不会被 MakeResident,GPU 第一次访问它就是非法访问。反过来,也正因为有这套机制,UE 才敢让总分配量超过显存预算:对象按需换入、按 LRU 换出,驻留量尽量压在预算以内。
三、跟踪谁:对象模型与粒度
3.1 三层对象
| 层级 | 类型 | 所有者 | 生命周期 |
|---|---|---|---|
| 被跟踪对象 | FD3D12ResidencyHandle,派生自 D3DX12Residency::ManagedObject | FD3D12Resource(committed)/ FD3D12Heap | 随资源 / heap |
| 单条 CL 的引用清单 | FD3D12ResidencySet = D3DX12Residency::ResidencySet | FD3D12CommandList,1:1 | 随 CL 对象,CL 池化复用 |
| LRU 与预算仲裁 | FD3D12ResidencyManager = D3DX12Residency::ResidencyManager | FD3D12Device 的内嵌成员 | 随 device |
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Residency.h:46
struct FD3D12ResidencyHandle : public D3DX12Residency::ManagedObject
{
#if DO_CHECK
class FD3D12GPUObject* GPUObject = nullptr; // 只用于 mGPU 下校验 GPUMask
#endif
};
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Residency.h:189
typedef D3DX12Residency::ResidencySet FD3D12ResidencySet;
typedef D3DX12Residency::ResidencyManager FD3D12ResidencyManager;
// Engine/Source/Runtime/D3D12RHI/Private/D3D12CommandList.h:96
FD3D12ResidencySet* const ResidencySet;
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Device.h:286
struct FResidencyManager : public FD3D12ResidencyManager
{
FResidencyManager(FD3D12Device& Parent);
~FResidencyManager();
} ResidencyManager;Manager 随 device 构造而初始化:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Device.cpp:166
FD3D12Device::FResidencyManager::FResidencyManager(FD3D12Device& Parent)
{
#if ENABLE_RESIDENCY_MANAGEMENT
IDXGIAdapter3* DxgiAdapter3 = nullptr;
VERIFYD3D12RESULT(Parent.GetParentAdapter()->GetAdapter()->QueryInterface(IID_PPV_ARGS(&DxgiAdapter3)));
const uint32 ResidencyMangerGPUIndex = GVirtualMGPU ? 0 : Parent.GetGPUIndex(); // GPU node index is used by residency manager to query budget
D3DX12Residency::InitializeResidencyManager(*this, Parent.GetDevice(), ResidencyMangerGPUIndex, DxgiAdapter3, RESIDENCY_PIPELINE_DEPTH);
#endif // ENABLE_RESIDENCY_MANAGEMENT
}RESIDENCY_PIPELINE_DEPTH = 6(D3D12RHIDefinitions.h:15),传进库里成为 MaxLatency,作用见第五节。
3.2 粒度:跟踪的是 ID3D12Pageable,不是逻辑资源
D3D12 只允许在 descriptor heap、heap、committed resource、query heap 上切换驻留状态:placed resource 借用的是 heap 的内存,reserved resource 自己没有物理内存。GetResidencyHandles() 正是按这个规则把资源映射到跟踪对象:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.h:400
TConstArrayView<FD3D12ResidencyHandle*> GetResidencyHandles() const
{
#if ENABLE_RESIDENCY_MANAGEMENT
if (!bRequiresResidencyTracking)
{
return {};
}
else if (IsPlacedResource())
{
return Heap->GetResidencyHandles(); // placed:所在的 heap
}
else if (IsReservedResource())
{
return ReservedResourceData->ResidencyHandles; // reserved:当前挂着的一组 backing heap
}
else
{
checkf(ResidencyHandle, TEXT("Resource requires residency tracking, but StartTrackingForResidency() was not called."));
return MakeArrayView(&ResidencyHandle, 1); // committed:资源本身
}
#else // ENABLE_RESIDENCY_MANAGEMENT
return {};
#endif // ENABLE_RESIDENCY_MANAGEMENT
}| 资源类型 | 被跟踪的 ID3D12Pageable | 注册入口 |
|---|---|---|
| Committed | 资源本身 | FD3D12Resource::StartTrackingForResidency(),D3D12Resources.cpp:625 |
| Placed | 所在的 heap | FD3D12Heap::BeginTrackingResidency(),D3D12Resources.cpp:757;各分配器的调用点在 D3D12PoolAllocator.cpp:110 / :507、D3D12Allocation.cpp:329、D3D12TransientResourceAllocator.cpp:83 |
| Reserved | 当前挂着的一组 backing heap,运行期会变 | CommitReservedResource() 新建 heap 时注册,并追加进 ReservedResourceData->ResidencyHandles,D3D12Resources.cpp:528-535 |
同一个 pool heap 里的所有 placed resource 共用一个跟踪对象:heap 里任何一个资源被登记,整个 heap 都会被换入。
3.3 哪些资源不跟踪
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp:142
#if ENABLE_RESIDENCY_MANAGEMENT
// Residency tracking is only used for GPU-only resources owned by the Engine.
// Back buffers may be referenced outside of command lists (during presents), however D3DX12Residency.h library
// uses fences tied to command lists to detect when it's safe to evict a resource, which is wrong for back buffers.
// External/shared resources may be referenced by command buffers in third-party code.
const bool bShouldTrackResource = (GNumExplicitGPUsForRendering == 1) || IsGPUOnly(InHeapType, HeapProps);
bRequiresResidencyTracking = !Desc.bExternal && !Desc.bBackBuffer && bShouldTrackResource;
#endif排除的是三类:
- Back buffer:present 时在 command list 之外被引用,而库判断”能否驱逐”靠的是绑在 CL 提交上的 fence,这个前提对 back buffer 不成立。
- External / shared 资源:可能被第三方代码的 command buffer 引用。
- 显式多卡下的非 GPU-only 资源:只有
GNumExplicitGPUsForRendering > 1时才要求 GPU-only;heap 侧的同一规则在FD3D12Heap::SetHeap()(D3D12Resources.cpp:712-715)。
❗ 注释第一句说只跟踪 GPU-only 资源,代码却在单卡时让 bShouldTrackResource 恒为 true:单卡下 upload / readback 资源同样被跟踪,创建时同样带 CREATE_NOT_RESIDENT。
StartTrackingForResidency() 对 back buffer 和 external 资源直接 checkf(D3D12Resources.cpp:634-635)。另有一个显式出口 FD3D12Heap::DisallowTrackingResidency()(D3D12Resources.cpp:749),头文件注释标明它是 UE-174791 / UE-202367 的 workaround 的一部分;它必须在 BeginTrackingResidency() 之前调用,否则 checkf 直接报错(D3D12Resources.cpp:752)。
四、一条 CL 的完整生命周期
4.1 调用链总览
【录制期 · RHI 线程 / 并行翻译线程】
FD3D12Device::ObtainCommandList() D3D12Device.cpp:703
├─ 池里没有 → new FD3D12CommandList() D3D12CommandList.cpp:98
│ ├─ ResidencySet = CreateResidencySet(...) D3D12CommandList.cpp:101 ← 创建(每个 CL 对象一次)
│ └─ D3DX12Residency::Open(ResidencySet) D3D12CommandList.cpp:189 ← 打开
└─ 池里有 → FD3D12CommandList::Reset() D3D12CommandList.cpp:204
└─ D3DX12Residency::Open(ResidencySet) D3D12CommandList.cpp:216 ← 重新打开
FD3D12ContextCommon::UpdateResidency(Resource) D3D12CommandContext.h:439
└─ FD3D12CommandList::UpdateResidency() D3D12CommandList.cpp:11
├─ 普通资源 → AddToResidencySet(handles) D3D12CommandList.cpp:41 ← 立即插入
└─ reserved → DeferredResidencyUpdateSet.Add() D3D12CommandList.cpp:16 ← 推迟到提交线程
FD3D12CommandList::Close() D3D12CommandList.cpp:224
└─ 无 deferred → D3DX12Residency::Close(ResidencySet) D3D12CommandList.cpp:248 ← 关闭(路径 A)
【提交期 · RHI Submission Thread】
FD3D12DynamicRHI::FlushBatchedPayloads() D3D12Submission.cpp:686
├─ payload 带 reserved commit → Flush() + UpdateReservedResources() D3D12Submission.cpp:887
│ └─ FD3D12Resource::CommitReservedResource() D3D12Resources.cpp:223
└─ Flush()
├─ ResidencySets.Add(CL->CloseResidencySet()) D3D12Submission.cpp:746
│ └─ 补登记 deferred 资源 + Close D3D12CommandList.cpp:26 ← 关闭(路径 B)
├─ FD3D12Queue::ExecuteCommandLists() WindowsD3D12Device.cpp:2468
│ └─ ResidencyManager.ExecuteCommandLists() WindowsD3D12Device.cpp:2481
│ └─ ExecuteSubset() d3dx12residency.h:1066 ← 合并、分页、提交、打 fence
└─ Device->ReleaseCommandList(CL) D3D12Submission.cpp:797 ← CL 回池
【销毁期】
~FD3D12CommandList() → DestroyResidencySet() D3D12CommandList.cpp:195
~FD3D12Resource() → EndTrackingObject() + delete 句柄 D3D12Resources.cpp:183
~FD3D12Heap() → EndTrackingObject() + delete 句柄 D3D12Resources.cpp:686
4.2 创建与打开
- set 在
FD3D12CommandList构造时创建(D3D12CommandList.cpp:101),Reset()不重建。CL 一提交完就被提交线程放回 device 的对象池(D3D12Submission.cpp:797→D3D12Device.cpp:721),下次ObtainCommandList()取出时走Reset()重新Open(),set 连同ppSet的内存一起复用。set 的内容在ExecuteSubset里已经合并进 MasterSet,所以立刻复用是安全的。 - residency 关闭时
CreateResidencySet返回nullptr(D3D12Residency.h:127-134),Open/Close/DestroyResidencySet的包装都会判空;AddToResidencySet里IsInitialized(Handle)恒为 false,根本走不到Insert。
Open() 的核心是抢一个并发槽位:
// Engine/Source/ThirdParty/Windows/D3DX12/Include/d3dx12residency.h:213
HRESULT Open()
{
Internal::ScopedLock Lock(&pSyncManager->MaskCriticalSection);
...
// Find the first available command list by bitscanning
for (UINT32 i = 0; i < ARRAYSIZE(pSyncManager->AvailableCommandLists); i++)
{
if (pSyncManager->AvailableCommandLists[i] == false)
{
CommandListIndex = i; // 本 set 的槽位号
pSyncManager->AvailableCommandLists[i] = true;
CommandlistAvailable = true;
break;
}
}
if (CommandlistAvailable == false)
{
// There are too many open residency sets, consider using less or increasing the value of MAX_NUM_CONCURRENT_CMD_LISTS
RESIDENCY_CHECK(false);
return E_OUTOFMEMORY;
}
CurrentSetSize = 0; // 逻辑清空,ppSet 的内存保留
IsOpen = true;
OutOfMemory = false;
return S_OK;
}CommandListIndex 是整套去重机制的钥匙。每个被跟踪对象身上都带着一张位图:
// Engine/Source/ThirdParty/Windows/D3DX12/Include/d3dx12residency.h:149
// This is used to track which open command lists this resource is currently used on.
bool CommandListsUsedOn[MAX_NUM_CONCURRENT_CMD_LISTS]; // MAX_NUM_CONCURRENT_CMD_LISTS = 1024(d3dx12residency.h:34)CommandListsUsedOn[CommandListIndex]一次下标访问就能判断对象是否已在本 set 里,去重是 O(1)。- 代价是每个跟踪对象固定多出 1KB(1024 个 bool)。跟踪对象只有 heap 和 committed resource,placed resource 不单独计;上万个跟踪对象就是十来 MB 的 CPU 内存。
- 槽位池属于
ResidencyManager(SyncManager的注释写着 “One per Residency Manager”,d3dx12residency.h:82),也就是每个FD3D12Device一个,容量 1024。任一时刻处于 Open 状态的 set 不能超过这个数,提交时临时建的 MasterSet 也要占一个。
4.3 插入:幂等去重
// Engine/Source/Runtime/D3D12RHI/Private/D3D12CommandList.cpp:11
void FD3D12CommandList::UpdateResidency(const FD3D12Resource* Resource)
{
#if ENABLE_RESIDENCY_MANAGEMENT
if (Resource->NeedsDeferredResidencyUpdate()) // 等价于 IsReservedResource()
{
State.DeferredResidencyUpdateSet.Add(Resource); // 推迟到提交线程
}
else
{
AddToResidencySet(Resource->GetResidencyHandles()); // 立即插入
}
#endif // ENABLE_RESIDENCY_MANAGEMENT
}
// Engine/Source/Runtime/D3D12RHI/Private/D3D12CommandList.cpp:41
void FD3D12CommandList::AddToResidencySet(TConstArrayView<FD3D12ResidencyHandle*> ResidencyHandles)
{
for (FD3D12ResidencyHandle* Handle : ResidencyHandles)
{
if (D3DX12Residency::IsInitialized(Handle))
{
#if DO_CHECK
check(Device->GetGPUMask() == Handle->GPUObject->GetGPUMask());
#endif
D3DX12Residency::Insert(*ResidencySet, *Handle);
}
}
}对外一般经 context 转发:FD3D12ContextCommon::UpdateResidency()(D3D12CommandContext.h:439)调用 GetCommandList().UpdateResidency(),而 GetCommandList() 在当前没有打开的 CL 时会先 OpenIfNotAlready()(D3D12CommandContext.h:335-340),所以经由 context 的登记总会落在一条 Open 状态的 CL 上。
库里的 Insert 靠位图去重:
// Engine/Source/ThirdParty/Windows/D3DX12/Include/d3dx12residency.h:184
inline bool Insert(ManagedObject* pObject)
{
RESIDENCY_CHECK(IsOpen);
RESIDENCY_CHECK(CommandListIndex != InvalidIndex);
// If we haven't seen this object on this command list mark it
if (pObject->CommandListsUsedOn[CommandListIndex] == false)
{
pObject->CommandListsUsedOn[CommandListIndex] = true;
if (ppSet == nullptr || CurrentSetSize >= MaxResidencySetSize)
{
Realloc(); // 首次 4096 项,之后按 1.5 倍扩容
}
if (ppSet == nullptr)
{
OutOfMemory = true;
return false;
}
ppSet[CurrentSetSize++] = pObject;
return true;
}
else
{
return false;
}
}同一对象在同一条 CL 里插 N 次只占一个位置,所以 descriptor cache 每次写描述符表都无条件登记一遍,开销也可以接受。
录制期的登记点分散在全模块一百二十处左右的 UpdateResidency 调用里,按类别归纳:
| 类别 | 代表位置 | 说明 |
|---|---|---|
| 资源屏障 | D3D12CommandList.cpp:71(FD3D12ContextCommon::AddBarrier) | 覆盖面最广:做过状态转换的资源自动入集合 |
| 描述符缓存 | D3D12DescriptorCache.cpp:SetVertexBuffers(:221)、BuildSRVTable(:461)、BuildUAVTable(:256)、SetConstantBufferViews(:571)、SetRootConstantBuffers(:756)、PrepareBindlessViews(:518 / :534)、SetRenderTargets(:312 / :324) | 绑定 VB、SRV / UAV / CBV 表、root CBV、bindless 视图和 RT / DS 时逐个登记 |
| RT / DS 清除 | D3D12Commands.cpp:1456 / 1464(RHIClearMRTImpl) | ClearRenderTargetView / ClearDepthStencilView 之后登记目标 |
| Copy / Update / Readback | D3D12Buffer.cpp、D3D12Texture.cpp、D3D12RenderTarget.cpp | src、dst 都要登记 |
| Pool 碎片整理拷贝 | D3D12PoolAllocator.cpp:1068-1069 | |
| Indirect args / UAV counter | D3D12Commands.cpp:115 / 1298、D3D12StateCache.cpp:1368 | 不经过描述符表,必须单独登记 |
| 光线追踪 | D3D12RayTracing.cpp,近 40 处 | BLAS / TLAS、scratch、instance buffer、shader table 以及 SBT 里引用的资源 |
| Query 结果回读 | D3D12Submission.cpp:569-574 | 提交线程上给临时取出的 resolve CL 直接 AddToResidencySet,登记 query heap 与结果 buffer |
| 对外显式接口 | ID3D12DynamicRHI::RHIUpdateResourceResidency(),实现在 D3D12RHI.cpp:831 | 给直接往 UE 的 D3D12 command list 里录命令的外部代码补登记 |
4.4 关闭:两条互斥路径
// Engine/Source/Runtime/D3D12RHI/Private/D3D12CommandList.cpp:246(FD3D12CommandList::Close 末尾,录制线程)
if (State.DeferredResidencyUpdateSet.Num() == 0)
{
D3DX12Residency::Close(ResidencySet); // 路径 A:没有 deferred 资源,录制线程上直接关
}
State.IsClosed = true;
// Engine/Source/Runtime/D3D12RHI/Private/D3D12CommandList.cpp:26(提交线程)
FD3D12ResidencySet* FD3D12CommandList::CloseResidencySet()
{
for (const FD3D12Resource* Resource : State.DeferredResidencyUpdateSet)
{
AddToResidencySet(Resource->GetResidencyHandles()); // 此刻才取 handle 快照
}
if (State.DeferredResidencyUpdateSet.Num() > 0)
{
D3DX12Residency::Close(ResidencySet); // 路径 B:有 deferred 资源,提交线程上关
}
return ResidencySet;
}两条路径以 DeferredResidencyUpdateSet 是否为空互斥,保证 set 恰好被 Close 一次。CloseResidencySet() 对每条 CL 都会调用(D3D12Submission.cpp:746),没有 deferred 资源时它只是把早已关好的 set 原样返回。走路径 B 的 CL,从录制线程关闭 D3D command list 到提交线程处理它的这段时间里,set 一直占着槽位。
4.5 消费:交给 Manager
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Submission.cpp:738(FlushBatchedPayloads 的 Flush lambda 内)
for (FD3D12CommandList* CommandList : Payload->CommandListsToExecute)
{
check(CommandList->IsClosed());
CommandList->FlushProfilerEvents(Payload->EventStream, Time);
D3DCommandLists.Add(CommandList->Interfaces.CommandList);
#if ENABLE_RESIDENCY_MANAGEMENT
ResidencySets.Add(CommandList->CloseResidencySet());
#endif
}
// Engine/Source/Runtime/D3D12RHI/Private/Windows/WindowsD3D12Device.cpp:2468
void FD3D12Queue::ExecuteCommandLists(TArrayView<ID3D12CommandList*> D3DCommandLists
#if ENABLE_RESIDENCY_MANAGEMENT
, TArrayView<FD3D12ResidencySet*> ResidencySets
#endif
)
{
TRACE_CPUPROFILER_EVENT_SCOPE(FD3D12Queue::ExecuteCommandLists);
#if ENABLE_RESIDENCY_MANAGEMENT
check(D3DCommandLists.Num() == ResidencySets.Num()); // 严格一一对应
if (GEnableResidencyManagement)
{
VERIFYD3D12RESULT(Device->GetResidencyManager().ExecuteCommandLists(
D3DCommandQueue,
D3DCommandLists.GetData(),
ResidencySets.GetData(),
D3DCommandLists.Num()
));
}
else
#endif
{
D3DCommandQueue->ExecuteCommandLists( // residency 关闭:裸提交
D3DCommandLists.Num(),
D3DCommandLists.GetData()
);
}
}一次 Flush() 攒下的 CL 会按 GetMaxExecuteBatchSize() 和 D3D12.MaxCommandsPerCommandList(默认 10000 条命令)切成若干批(D3D12Submission.cpp:758-777),每批对应一次 ResidencyManager::ExecuteCommandLists,也就是一次合并、分页和 fence 打点。
4.6 销毁
~FD3D12CommandList()里DestroyResidencySet()(D3D12CommandList.cpp:195)。CL 用完是回池而不是析构,set 跟着 CL 对象长期存活。- 对象侧的注销在
~FD3D12Resource()(D3D12Resources.cpp:180-186)和~FD3D12Heap()(D3D12Resources.cpp:683-689):EndTrackingObject()把自己从 LRU 链表摘掉,再delete句柄。
五、提交那一刻:ExecuteSubset 与 ProcessPagingWork
ResidencyManager::ExecuteCommandLists() 直接转到 ExecuteSubset()(d3dx12residency.h:1066-1189):
1. 校验:任一 set 仍处于 Open → 返回 E_INVALIDARG
// Residency Sets must be closed before execution just like Command Lists
2. 建一个临时 MasterSet,把这一批所有 set 合并去重
TotalSizeNeeded 累加每个唯一对象的 Size(已经 resident 的也算)
MasterSet 同样要 Open,合并完立刻 Close 归还槽位
3. if (Count > 1 && TotalSizeNeeded > Local.Budget + NonLocal.Budget)
→ 对半拆开,递归 ExecuteSubset(前一半) + ExecuteSubset(后一半)
Count == 1 时拆无可拆,照常往下走
4. EnqueueAsyncWork(MasterSet) → ProcessPagingWork() d3dx12residency.h:1257
换入、驱逐都在这里完成
5. AsyncThreadFence.GPUWait(Queue) ← 队列在 GPU 侧等换入完成
6. Queue->ExecuteCommandLists(Count, CommandLists) ← 真正提交
7. SignalFence(Queue, QueueFence) ← 打一个 device 级 sync point,LRU 据此判断 GPU 是否用完
第 5 步是整套机制的安全闸:队列在 GPU 侧等待换入完成的 fence,GPU 不会在资源真正 resident 之前开始执行这批 CL。
5.1 ProcessPagingWork
ProcessPagingWork()(d3dx12residency.h:1257-1477)按顺序做四件事:
- 标记:MasterSet 里 EVICTED 的对象进换入列表,并在 LRU 里转为 RESIDENT;所有对象更新
LastGPUSyncPoint(本批 sync point 编号)和LastUsedTimestamp(当前 QPC),挪到 LRU 链表尾部。 - 按闲置时间驱逐:
TrimAgedAllocations()(d3dx12residency.h:627)从链表头(最久未用)开始,驱逐”最后一次使用的 sync point 已在 GPU 上完成、且闲置超过 grace period”的对象,批量Device->Evict()。grace period 的算法见第六节。 - 按预算换入:查询 Local + NonLocal 的 Budget 与 CurrentUsage,按剩余空间分批换入;支持
ID3D12Device3时走EnqueueMakeResident,换入完成由 D3D signal 分页 fence。 - 空间不够时腾地方:CPU 阻塞等待最早一个未完成的 sync point(
WaitForSyncPoint),再用TrimToSyncPointInclusive()(d3dx12residency.h:601)按 LRU 顺序驱逐”最后使用于该 sync point 及之前”的对象,直到用量回到预算以内,然后继续换入。
实在无对象可驱逐时(LRU 头部就是本批要用的对象,或者已经没有在途的 sync point),库会无视预算把剩下的对象全部换入(d3dx12residency.h:1399-1435)。这一步若再失败,源码里只留了一句 TODO(“catastrophic failure”),错误码不会返回给调用方,而这些对象在第 1 步已经被 LRU 记成 RESIDENT,后续提交也不会再尝试换入。
5.2 两种执行模式
支持 ID3D12Device3 | 不支持 | |
|---|---|---|
ProcessPagingWork 在哪跑 | 调用 ExecuteCommandLists 的线程,即 RHI 提交线程 | 库自己的异步分页线程,只在不支持 Device3 时创建(d3dx12residency.h:772-782) |
| 换入 API | EnqueueMakeResident,完成时由 D3D signal fence | 阻塞的 MakeResident,返回后由 CPU signal fence |
MaxLatency(= 6) | 工作队列从不积压,基本不起作用 | 提交线程最多领先分页线程 6 个 workload,超出就等(d3dx12residency.h:1482-1486) |
走默认 EVICTED 路径的机器必然拿得到 ID3D12Device8,也就必然支持 ID3D12Device3,所以常见情况是左列:第 4 步的”CPU 阻塞等 GPU”就发生在 RHI 提交线程上,超预算时表现为提交线程卡顿。
5.3 单条 CL 超预算
第 3 步的拆分以 CL 为最小单位。单条 CL 引用的对象总量本身就超过预算时拆无可拆,只能走上面的”无视预算强行换入”;要么减少单条 CL 的引用量,要么降低整体显存压力。
六、移除到底发生在哪一层
ResidencySet 没有针对单个对象的删除接口,Close() 里的 Remove() 名字很有误导性:
// Engine/Source/ThirdParty/Windows/D3DX12/Include/d3dx12residency.h:252
HRESULT Close()
{
if (IsOpen == false)
{
return E_INVALIDARG;
}
if (OutOfMemory == true)
{
return E_OUTOFMEMORY;
}
for (INT32 i = 0; i < CurrentSetSize; i++)
{
Remove(ppSet[i]);
}
ReturnCommandListReservation();
IsOpen = false;
return S_OK;
}
// d3dx12residency.h:278
inline void Remove(ManagedObject* pObject)
{
pObject->CommandListsUsedOn[CommandListIndex] = false; // 只清去重标记
}
// d3dx12residency.h:283
inline void ReturnCommandListReservation()
{
Internal::ScopedLock Lock(&pSyncManager->MaskCriticalSection);
pSyncManager->AvailableCommandLists[CommandListIndex] = false; // 归还槽位
CommandListIndex = ResidencySet::InvalidIndex;
IsOpen = false;
}Close() 清掉的只是每个对象上的去重标记,然后把槽位还回池子;ppSet[] 的内容原封不动,提交时要读的正是它。
“移除”其实分三层:
| 层级 | 何时”移除” | 机制 |
|---|---|---|
| 集合层 | 对象插入后一直留到提交;下一次 Open() 把 CurrentSetSize 归零,整体作废 | 逻辑清零,ppSet 内存复用 |
| 驻留层 | 由 Manager 的 LRU 决定 | TrimAgedAllocations() / TrimToSyncPointInclusive() → Device->Evict() |
| 注册层 | 资源或 heap 析构 | EndTrackingObject() 摘出 LRU,delete 句柄 |
6.1 驻留层:什么时候真正换出
两条路径都在 ProcessPagingWork() 里(第五节第 2、4 步):一条按闲置时间,每次提交都会跑;一条按预算,只在换入放不下时触发。按闲置时间驱逐的门槛由 grace period 决定:
// Engine/Source/ThirdParty/Windows/D3DX12/Include/d3dx12residency.h:1625
UINT64 GetCurrentEvictionGracePeriod(DXGI_QUERY_VIDEO_MEMORY_INFO* LocalMemoryState)
{
// 1 == full pressure, 0 == no pressure
double Pressure = (double(LocalMemoryState->CurrentUsage) / double(LocalMemoryState->Budget));
Pressure = RESIDENCY_MIN(Pressure, 1.0);
if (Pressure > cTrimPercentageMemoryUsageThreshold) // 0.7
{
// Normalize the pressure for the range 0 to cTrimPercentageMemoryUsageThreshold
Pressure = (Pressure - cTrimPercentageMemoryUsageThreshold) / (1.0 - cTrimPercentageMemoryUsageThreshold);
// Linearly interpolate between the min period and the max period based on the pressure
return UINT64((MaxEvictionGracePeriodTicks - MinEvictionGracePeriodTicks) * (1.0 - Pressure)) + MinEvictionGracePeriodTicks;
}
else
{
// Essentially don't trim at all
return MAXUINT64;
}
}| 本地显存 CurrentUsage / Budget | grace period |
|---|---|
≤ 70%(cTrimPercentageMemoryUsageThreshold) | MAXUINT64,完全不按时间驱逐 |
| 70% ~ 100% | 从 60s 线性收缩到 1s(cMaxEvictionGracePeriod / cMinEvictionGracePeriod,d3dx12residency.h:688-689) |
LocalMemoryBudgetLimit == 0 | 直接归零(d3dx12residency.h:1312-1317);按预算驱逐时的目标预算也按 0 算(d3dx12residency.h:1449-1453) |
LocalMemoryBudgetLimit 是 Epic 加进库里的本地预算上限,GetCurrentBudget() 取它与 DXGI Budget 的较小值(d3dx12residency.h:1530-1535)。它在 FD3D12Adapter::CollectMemoryStats() 里更新(D3D12Adapter.cpp:1920-1939):
- 正常情况:DXGI 报告的本地 Budget;
D3D12.EvictAllResidentResourcesInBackground打开且窗口失焦:0。源码注释给的理由是,多个程序同时大量占用显存时 DXGI 报告的 Budget 依然偏高,VidMM 会按自己的启发式把分配挪出显存,被挪走的资源在挪回来之前被拿去渲染,性能会明显下降;不如失焦时主动把预算压成 0,立刻腾出显存;D3D12.ResidencyDebugBudgetMB≥ 0:用调试值覆盖,见第九节。
七、Reserved Resource 的延迟登记
7.1 问题:录制时不知道它挂着哪些 heap
// Engine/Source/Runtime/D3D12RHI/Private/D3D12CommandList.h:300
// Resources whose residency must be updated on the submission thread, as their residency handles are not known during translation.
// This includes reserved resources that may refer to different heaps at different points on the submission timeline.
TSet<const FD3D12Resource*> DeferredResidencyUpdateSet;
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.h:363
inline bool NeedsDeferredResidencyUpdate() const { return IsReservedResource(); }Reserved resource 的 backing heap 由 CommitReservedResource() 在提交线程上增删,而 commit 请求是录制期排进 payload 的:
【录制线程】
FD3D12ContextCommon::SetReservedBufferCommitSize() D3D12CommandContext.cpp:332
├─ if (IsPendingCommands()) CloseCommandList() ← 收尾当前 CL(它在 Open 时已挂进当前 payload 的 Execute 阶段)
└─ GetPayload(EPhase::UpdateReservedResources)
->ReservedResourcesToCommit.Add(CommitDesc) D3D12CommandContext.cpp:345
【提交线程】FlushBatchedPayloads() D3D12Submission.cpp:686
for each payload:
if (Payload->HasUpdateReservedResourcesWork()) D3D12Submission.cpp:887
{
Flush(); ← 先执行之前攒下的 CL(此时取它们的 handle 快照)
UpdateReservedResources(Payload); ← 再改 tile 映射,backing heap 在这里创建 / 回收
}
...
Flush() → CloseResidencySet() → ExecuteCommandLists() ← 本 payload 的 CL 在 commit 之后才执行
GetPayload() 发现当前阶段已经越过 UpdateReservedResources,就会新开一个 payload(D3D12CommandContext.h:354-363),commit 于是被夹在前后两批 CL 之间:之前录制的 CL 看到的是 commit 前的 heap 组合,之后录制的看到的是 commit 后的。同一个 reserved 资源,在提交时间线的不同位置对应不同的 heap 列表,而录制线程(包括并行翻译线程)无从知道一条 CL 最终排在哪次 commit 之后。
7.2 解法:只记资源指针,提交前再取快照
录制期只把 FD3D12Resource* 放进 DeferredResidencyUpdateSet;提交线程在 Flush() 里对每条 CL 调 CloseResidencySet(),这时才用 GetResidencyHandles() 取当下的 heap 列表插进 set。FlushBatchedPayloads() 在执行 commit 之前总会先 Flush() 掉更早的 CL(D3D12Submission.cpp:887-891),所以每条 CL 取到的快照都恰好对应它在时间线上的位置。这也是 FD3D12CommandList::Close() 在 deferred 集合非空时不关 set 的原因:set 要留给提交线程补插。
同一个原因让 reserved resource 的驻留状态在 RHI 线程上无从查询:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.h:367
bool IsResident() const
{
#if ENABLE_RESIDENCY_MANAGEMENT
if (NeedsDeferredResidencyUpdate())
{
// We don't know the state because the set of residency handles is only known on the
// RHI Submission Thread and may change throughout the frame.
return true;
}
...7.3 Commit 本身也要用一个 ResidencySet
UpdateTileMappings 是 queue 级操作,不属于任何 CL,但新 heap 同样必须先 resident 才能被映射过去。CommitReservedResource() 为此临时建一个不绑定 CL 的 set(D3D12Resources.cpp:553-579),交给 Epic 在库里加的 ResidencyManager::MakeResident(Queue, Set)(d3dx12residency.h:935-987,注释写明是为 UpdateTileMappings 这类 queue 级操作准备的,逻辑复制自提交路径);映射完成后再 SignalFence(Queue)(D3D12Resources.cpp:594-599),相当于告诉 LRU”这些 heap 被这个队列用到了这个时间点”。Commit 的完整流程见 UE5 DX12 Reserved Resource 从硬件语义到引擎落地。
八、开关与常量
| CVar / 宏 | 默认值 | 位置 | 作用 |
|---|---|---|---|
ENABLE_RESIDENCY_MANAGEMENT | Windows 为 1 | D3D12RHI.h:29 | 编译期总开关。其他平台由各自的 D3D12RHIPlatformPublic.h 决定;平台把 D3D12_PLATFORM_NEEDS_RESIDENCY_MANAGEMENT 定为 0 时,D3D12Residency.h:15-16 用 static_assert 强制它也为 0 |
D3D12.ResidencyManagement | 1,ECVF_ReadOnly | D3D12Adapter.cpp:31 | 建 device 时读取(D3D12Adapter.cpp:771),为 0 则 GEnableResidencyManagement = false:包装函数全部空转,CreateResidencySet 返回 nullptr,提交走裸 ExecuteCommandLists。单独关它不够,见第九节 |
D3D12.ResourcesStartResident | 0,ECVF_ReadOnly | D3D12Adapter.cpp:39 | 0:跟踪对象 EVICTED 出生(heap 带 CREATE_NOT_RESIDENT);1:创建即 RESIDENT |
D3D12.EvictAllResidentResourcesInBackground | 0 | D3D12Adapter.cpp:197 | 窗口失焦时把本地预算压成 0,尽可能驱逐所有闲置的跟踪对象 |
D3D12.ResidencyDebugBudgetMB | -1(关闭) | D3D12Adapter.cpp:204 | ≥ 0 时覆盖本地预算,人为制造分页压力 |
D3D12.MaxCommandsPerCommandList | 10000 | D3D12CommandContext.cpp:15 | 单条 CL 命令数超过它就切 CL;提交时也按它把 CL 分批送进 ExecuteCommandLists |
RESIDENCY_PIPELINE_DEPTH | 6 | D3D12RHIDefinitions.h:15 | Manager 的 MaxLatency,限制提交线程领先异步分页线程的 workload 数 |
MAX_NUM_CONCURRENT_CMD_LISTS | 1024 | d3dx12residency.h:34 | 每个 Manager 同时处于 Open 的 set 上限 |
RESIDENCY_SINGLE_THREADED | 0 | d3dx12residency.h:28 | 0 时允许异步分页线程;设备支持 ID3D12Device3 时库自动走单线程 |
| grace period | 1s ~ 60s | d3dx12residency.h:688-689 | 本地用量超过预算 70% 后随压力线性收缩 |
九、踩坑与排查
9.1 写代码时
-
新增直接录制 GPU 命令的路径,要配套
UpdateResidency。 走AddBarrier或 descriptor cache 的资源会被自动覆盖;indirect args buffer、UAV counter、query heap、光追 shader table 里的间接引用都要手动补。在 UE 之外直接往 UE 的 command list 里录命令的代码,用ID3D12DynamicRHI::RHIUpdateResourceResidency()(ID3D12DynamicRHI.h:65)登记。 -
❗
Insert只能发生在 set 打开期间,而且没有断言兜底。RESIDENCY_CHECK在库里默认是空宏(d3dx12residency.h:16-25,只有#if 0分支才会DebugBreak),Insert开头那两句检查在 UE 里不生效。往已经 Close 的 set 里插入时CommandListIndex已是InvalidIndex(0xFFFFFFFF),CommandListsUsedOn[CommandListIndex]直接越界。经由FD3D12ContextCommon的登记碰不上这个问题;在提交线程上拿FD3D12CommandList*直接AddToResidencySet时要自己保证 set 还开着,query resolve 就是专门现取一条新 CL 来登记(D3D12Submission.cpp:554-563)。 -
mGPU 下 GPUMask 必须匹配。
AddToResidencySet里的check(Device->GetGPUMask() == Handle->GPUObject->GetGPUMask())(D3D12CommandList.cpp:48,仅DO_CHECK构建)会拦住跨 device 登记的 handle。 -
别把 back buffer / external 资源接进跟踪。 库的”可驱逐”判断建立在 CL 提交的 fence 上,这两类资源会在 CL 之外被访问;
StartTrackingForResidency()对它们直接checkf。 -
控制单条 CL 的引用量。 预算不够时的拆分以 CL 为最小单位,单条 CL 引用量本身超预算,只能越过预算强行换入;这一步再失败,错误码被库吞掉,GPU 随后访问的仍是未驻留对象。
-
CL 切得越碎,固定开销越明显。 每个 set 的
Open要进临界区线性扫描槽位,Close要逐个清掉对象上的去重标记;每批ExecuteCommandLists还要合并一次 MasterSet、打一次 sync point。
9.2 排查时
-
漏登记在显存充裕时很难暴露。 驻留单位是 heap,同一个 pool heap 里别的资源被登记过,整个 heap 就是 resident 的;本地显存用量不到预算的 70% 时 LRU 不按时间驱逐,对象只要被某条 CL 换入过就会一直留着。漏登记通常要等显存吃紧、对象真被换出之后,才以
DEVICE_REMOVED的形式出现。 -
用
D3D12.ResidencyDebugBudgetMB把问题逼出来。 帮助文本写明了它的用途:暴露那些平时被驱动隐式分页掩盖的漏登记。设一个很小的值(帮助文本建议 128,0 为最大压力),驱逐和换入会变得极其频繁,漏登记在开发期就能稳定复现。 -
❗ 关闭 residency 管理,要两个 CVar 一起设。
D3D12.ResidencyManagement=0能把问题范围快速收敛到”是不是漏跟踪”,但它只让GEnableResidencyManagement = false;heap 带不带CREATE_NOT_RESIDENT只看GD3D12StartResourceResident。只关前者,heap 照样以未驻留状态创建,却再也没有人MakeResident。两个都是ECVF_ReadOnly,放在 ini 的[SystemSettings]里:[SystemSettings] D3D12.ResidencyManagement=0 D3D12.ResourcesStartResident=1非 Shipping 包也可以走命令行覆盖 ini:
-ini:Engine:[SystemSettings]:D3D12.ResidencyManagement=0,[SystemSettings]:D3D12.ResourcesStartResident=1这样所有对象常驻、不再驱逐,显存占用会明显上升,只适合排查,不适合出包。
-
这里报出的 “Out of video memory” 不一定是显存不够。
Open/Close/ExecuteCommandLists的返回值都要过VERIFYD3D12RESULT(D3D12Residency.h:151/:161、WindowsD3D12Device.cpp:2481),而VerifyD3D12Result遇到E_OUTOFMEMORY一律转进TerminateOnOutOfMemory()(D3D12Util.cpp:1070-1077),弹出 “Out of video memory trying to allocate a rendering resource…” 后退出。residency 库在 1024 个槽位用光、ppSet扩容失败、MasterSet 分配失败时返回的都是E_OUTOFMEMORY。ResidencySet::Initialize里有一段 Epic 改动专门放过空 set(d3dx12residency.h:304-317),注释描述的正是这种以”驱动崩溃”提示收场的误报。 -
超预算时的卡顿看提交线程。 支持
ID3D12Device3时ProcessPagingWork跑在 RHI 提交线程上,腾空间要 CPU 阻塞等 GPU 完成更早的 sync point。Insights 里FD3D12Queue::ExecuteCommandLists(WindowsD3D12Device.cpp:2474的 CPU 事件)明显变长时,先看显存是否已经超预算。
十、总结
ResidencySet 把”某条 command list 的显存工作集”这个原本只有 GPU 执行时才确定的信息,提前物化成 CPU 侧一份 O(1) 去重的清单。有了它,ResidencyManager 才能在提交那一刻做三件事:
- 把要用的换进来(
MakeResident/EnqueueMakeResident,并让队列在 GPU 侧等换入完成); - 把 GPU 已经用完、闲置超过 grace period 的换出去(LRU trim +
Evict); - 必要时把一批 CL 对半拆开,塞进当前预算。
这正是 UE 在 Windows 上敢让跟踪对象默认 EVICTED 出生、把总分配量放到显存预算之外的基础。反过来说,任何绕过 barrier 和 descriptor cache 直接触碰资源的新路径,都得记得给它登记一笔。
十一、参考链接
🔗 Microsoft Learn - Residency(D3D12)
🔗 Microsoft Learn - Direct3D 12 residency starter library
🔗 GitHub - DirectX-Graphics-Samples:d3dx12Residency.h
🔗 Microsoft Learn - D3D12_HEAP_FLAGS(D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT)
🔗 Microsoft Learn - ID3D12Device3::EnqueueMakeResident
🔗 Microsoft Learn - ID3D12Device::Evict