来源:DX12 引擎源码
引擎版本:UE 5.8
面向读者:UE 图形程序员
DX12 里 committed / placed / reserved 三种资源,前两种在 UE 里天天见,reserved 却直到 5.8 才在 PC 上默认跑起来 —— GPUScene 的四个大 buffer、Nanite 的 ClusterPageData、VSM 的 physical page pool,全都是它的消费者。
有意思的是:UE 用它,用的几乎都不是它最出名的那个能力(稀疏驻留)。
目录
- 三种资源,同一张坐标系
- 硬件与 API 层原理
- UE5 的接入面
- 核心实现精读:CommitReservedResource
- Buffer 与 Texture 的分野
- UE5 里的三个消费者
- VA 账单与调优旋钮
- 收益与代价清单
- 如何验证它在工作
- 总结
- 参考链接
一、三种资源,同一张坐标系
DX12 资源 = GPU 虚拟地址空间(VA) + 物理页(heap) + 两者之间的映射(页表)。DX11 把这三者绑死成一个东西,DX12 才把它们拆开:
| VA 谁分配 | 物理内存谁分配 | 映射关系 | 能否改映射 | |
|---|---|---|---|---|
| Committed | 运行时隐式 | 运行时隐式(隐式 heap,刚好装下资源) | 创建时 1:1 固定 | ❌ |
| Placed | 运行时(落在 heap 的 VA 里) | 先 CreateHeap,再给 offset | 创建时固定,必须连续 | ❌(要换只能销毁重建 + 重建 descriptor) |
| Reserved | 只分配 VA,全部页初始为 NULL | 完全由你控制,可来自多个 heap | 运行时任意改,可稀疏 | ✅ UpdateTileMappings |
❗ 一个命名陷阱先说在前面:UE 里 CommitReservedResource() 的 “commit”,指的是给 reserved 资源的一段虚拟范围建立物理 backing,不是把它变成上表里的 Committed Resource。资源从 CreateReservedResource() 创建的那一刻起就永远是 reserved,不存在「转化」这回事。
微软对 reserved 的定义是:它拥有自己独立的 GPU 虚拟地址空间,允许先大量预留 VA,之后再把 VA 页映射到 heap 的某些区域,并且随时可以重新配置;当一段 VA 映射到 NULL 或未驻留的 heap 时,该部分资源被视为 non-resident。
PIX 团队给过一个很好的一句话类比:placed resource 在概念上就是一个只有一条、且永远不能改的映射的 reserved resource。
跨 API 对照:Reserved 就是 D3D11 的 Tiled Resources,也就是 Vulkan 的 Sparse Binding / Sparse Residency、OpenGL 的 Partially Resident Textures。
二、硬件与 API 层原理
2.1 Tile:映射的最小单位是 64KB
Reserved Resource 的 VA 空间(连续、巨大、免费)
┌──────┬──────┬──────┬──────┬──────┬──────┬──────┬──────┐
│ T0 │ T1 │ T2 │ T3 │ T4 │ T5 │ T6 │ T7 │ 每格 = 64KB tile
└──┬───┴──┬───┴──┬───┴──┬───┴──────┴──────┴──────┴──────┘
│ │ │ │ ↑ 这些 tile 映射到 NULL(不占物理显存)
▼ ▼ ▼ ▼
┌────────────────┐ ┌────────────────┐
│ ID3D12Heap A │ │ ID3D12Heap B │ ← 物理页可来自多个 heap(D3D12 相对 D3D11 的关键放宽)
└────────────────┘ └────────────────┘
- Buffer:64KB 线性切分,
GetResourceTiling返回一个平凡的 1D subresource tiling。 - Texture:tile 形状是「各维度近似等长的 64KB 区域」,这是
D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE的定义要求。
3D(Volume)的标准 tile 形状官方给了表:8bpp→64×32×32、16bpp→32×32×32、32bpp→32×32×16、64bpp→32×16×16、128bpp→16×16×16、BC1/4→128×64×16、BC2/3/5/6/7→64×64×16。
2D 情形同理由 64KB 约束推出:8bpp 256×256、16bpp 256×128、32bpp 128×128、64bpp 128×64、128bpp 64×64 —— 每项乘起来正好 65536 字节。
⚠️ 这组 2D 数值没有对应的官方表页面,是从「64KB + 近似等边」两条约束反推的业界通行值。工程上不要硬编码,运行时以 GetResourceTiling 返回的 D3D12_TILE_SHAPE 为准(UE 就是这么做的)。
分层视角:一次访问要穿过五层
上面那张图只画了「VA tile ↔ heap tile」这一层。完整链路有五层,混淆它们是理解 reserved resource 时最常见的卡点:
Reserved resource 的分层内存模型:tile mapping 层与 residency 层是两件独立的事
关键在于 tile mapping 层和 residency 层相互独立:一个 resource tile 已经指向某个 heap tile(映射建好了),不代表那个 heap 此刻驻留在 GPU 可访问的内存里。UE 分别用 UpdateTileMappings 和 Residency Manager 管这两层,两者之间的时序约束见 §4.4 阶段 6。
2.2 Packed Mip(mip tail)—— 最容易踩坑的地方
当某级 mip 的尺寸小于一个 tile 时,硬件会把剩下的所有小 mip 打包进 1~N 个 tile,称为 packed mips / mip tail。GetResourceTiling 通过 D3D12_PACKED_MIP_INFO 告诉你 NumStandardMips、NumPackedMips、NumTilesForPackedMips。
约束是硬的:
- packed mip 必须整体一次性映射完,不能只映射其中一部分(UE 代码里对此有专门的 assert,见 §4.5);
- Tier 2/3 下,「数组切片 > 1」且「存在小于 tile 的 mip」的组合不允许;Tier 4 才放开这条限制。
2.3 Tier 分级:决定你能不能用、用得多放心
| Tier | 能力 | 说明 |
|---|---|---|
| NOT_SUPPORTED | CreateReservedResource 完全不可用,连 buffer 都不行 | |
| Tier 1 | 可创建 2D reserved 纹理与 buffer | ⚠️ 读写 NULL tile 是 undefined;无 LOD clamp 指令;官方建议把同一个 page 重复映射到所有「本该 NULL」的位置来规避 |
| Tier 2 | mip 组织有明确保证;读 NULL tile 返回 0,写 NULL tile 被丢弃;shader 提供 LOD clamp 与 residency 反馈(Sample(S,float,int,float,uint) + CheckAccessFullyMapped) | Feature Level 12_0 的适配器全部保证 ≥ Tier 2 |
| Tier 3 | 增加 3D(Volume)Tiled Resources | |
| Tier 4 | 数组纹理可带完整 mip 链(含 packed mip) |
2.4 API 面:比 placed 多出来的那一套
| API | 作用 | 执行位置 |
|---|---|---|
ID3D12Device::CreateReservedResource(1) | 只建 VA,无 backing store | 立即(CPU) |
ID3D12Device::GetResourceTiling | 查询 tile 数、tile 形状、packed mip 信息 | 立即(CPU) |
ID3D12CommandQueue::UpdateTileMappings | 改页表:把 tile 区间映射到 heap 区间 / NULL / 单 tile 复用 / 跳过 | Queue 级操作,不是 command list 操作 |
ID3D12CommandQueue::CopyTileMappings | 把一个 reserved 资源的映射整体拷到另一个 | Queue 级 |
ID3D12GraphicsCommandList::CopyTiles | 在 tile 与线性 buffer 之间搬数据(不是映射) | Command list |
UpdateTileMappings 的 range flag 四选一:NONE(顺序映射一段 heap tile)、REUSE_SINGLE_TILE(多个 VA tile 共享同一物理 tile,这就是 Tier 1 规避 NULL 的手段)、NULL(解映射)、SKIP(保持原样)。
❗「Queue 级」这三个字是理解 UE 实现的钥匙:它按队列提交顺序执行,不进 command list,所以引擎必须把它安排在正确的提交时点上,并且跨队列时要显式 fence 同步(见 §3.3、§3.4)。
2.5 NVIDIA 的实战建议
- 把所有
UpdateTileMappings移到异步 copy queue,用来隐藏 OS 调度与提交开销; - 只要是重新映射同一批 tile,就不需要显式 unmap;
- 避免用
pHeap = NULL/D3D12_TILE_RANGE_FLAG_NULL做显式 unmap —— 那会强制驱动遍历所有映射并移除不再映射的 tile;宁可切到另一个 tile; - 按 2MB 粒度对齐 tile 更新,以便获得最优的硬件压缩(DCC 等)。
其中第一条和 UE 的选择是相反方向的取舍,见 §3.4。
三、UE5 的接入面
3.1 能力检测:门槛定在 Tier 2
// Engine/Source/Runtime/D3D12RHI/Private/Windows/WindowsD3D12Device.cpp:1904
// Tier 2 is guaranteed for all adapters with feature level 12_0.
GRHIGlobals.ReservedResources.Supported = Options.TiledResourcesTier >= D3D12_TILED_RESOURCES_TIER_2;
// Tier 3 is required to create volume textures. Some hardware may support it.
GRHIGlobals.ReservedResources.SupportsVolumeTextures = Options.TiledResourcesTier >= D3D12_TILED_RESOURCES_TIER_3;UE 把门槛直接定在 Tier 2,而不是 API 最低要求的 Tier 1 —— 因为 Tier 1 的「读 NULL tile 未定义」对引擎来说没法安全使用。
RHI 全局能力块在 RHIGlobals.h:720-751:
struct FReservedResources
{
bool Supported = false;
bool SupportsVolumeTextures = false;
int32 TextureArrayMinimumMipDimension = 256; // 规避 §2.2 的 array + 小 mip 限制(保守、与格式无关)
static constexpr int32 TileSizeInBytes = 65536; // 跨平台统一为 64KB
volatile int64 VirtualSize = 0; // 所有 reserved 资源占用的 VA 总量
} ReservedResources;TextureArrayMinimumMipDimension = 256 这个保守值,就是为了不去碰「Tier 4 才放开的那条限制」。
3.2 创建侧:两个新 flag 与几条硬约束
RHIDefinitions.h:966 / :1171:
EBufferUsageFlags::ReservedResource(BUF_ReservedResource)ETextureCreateFlags::ReservedResource(TexCreate_ReservedResource)ETextureCreateFlags::ImmediateCommit(TexCreate_ImmediateCommit)—— 创建即全量提交
注释里明确写了 EXPERIMENTAL,且不可与 Dynamic 等「不能落在 local GPU memory」的 flag 共用。
创建路径 D3D12Resources.cpp:887 里的硬性 check 正好对应 §2.1。注意 layout 那条只对 1D/2D/3D 纹理生效,buffer 不受此约束:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp
if (LocalDesc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE1D
|| LocalDesc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE2D
|| LocalDesc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE3D)
{
checkf(LocalDesc.Layout == D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE, ...);
}
checkf(LocalDesc.Alignment == 0 || LocalDesc.Alignment == 65536, ...);
// ...
// NOTE: reserved resource residency is not tracked/managed by the engine, so we don't need to call StartTrackingForResidency().最后那行 NOTE 很重要:reserved 资源本身不做 residency 追踪,驻留单位是 backing heap(见 §4.4 阶段 6)。
Buffer 侧的约束在 D3D12Buffer.cpp:486:
checkf(!bHasInitialData, TEXT("Reserved resources may not have initial data"));
checkf(!bIsDynamic, TEXT("Reserved resources may not be dynamic"));
checkf(!ResourceAllocator, TEXT("Reserved resources may not use a custom resource allocator"));3.3 提交侧:UE 把 commit 挂在了 Transition 上
这是一个很「UE」的设计选择 —— 不新增一条 RHI 命令,而是复用资源转换:
// Engine/Source/Runtime/RHI/Public/RHITransition.h:103-116
/**
* Represents a change in physical memory allocation for a resource that was created with TexCreate/BUF_ReservedResource flag.
* Physical memory is allocated in tiles/pages and mapped to the tail of the currently committed region of the resource.
* This API may be used to grow or shrink reserved resources without moving the bulk of the data or re-creating SRVs/UAVs.
* The contents of the newly committed region of the resource is undefined and must be overwritten by the application before use.
* Reserved resources must be created with maximum expected size, which will not cost any memory until committed.
* Commit size must be smaller or equal to the maximum resource size specified at creation.
* Check GRHIGlobals.ReservedResources.Supported before using this API or TexCreate/BUF_ReservedResource flag.
*/
struct FRHICommitResourceInfo
{
uint64 SizeInBytes = 0;
explicit FRHICommitResourceInfo(uint64 InSizeInBytes) : SizeInBytes(InSizeInBytes) {}
};调用链(自上而下):
FRDGBuilder::QueueCommitReservedBuffer(Buffer, NewSize) RenderGraphBuilder.inl:427
└─> FRDGPooledBuffer::SetCommittedSize() RenderGraphResources.h:1273
└─> FRHITransitionInfo{ Buffer, Before, After, FRHICommitResourceInfo }
└─> FD3D12ContextCommon::SetReservedBufferCommitSize() D3D12CommandContext.cpp:332
├─ if (IsPendingCommands()) CloseCommandList() ← 先关掉当前 command list
└─ GetPayload(EPhase::UpdateReservedResources)->ReservedResourcesToCommit.Add(...)
└─> FD3D12DynamicRHI::UpdateReservedResources(Payload) D3D12Submission.cpp:649
└─> FD3D12Resource::CommitReservedResource(Queue, Size) D3D12Resources.cpp:223
这里的 CloseCommandList() 和独立的 EPhase::UpdateReservedResources payload 阶段,正是 §2.4 那句「queue 级操作」在架构上的落地:页表更新必须切在两批 command list 之间,不能混在里面。
而且切断发生在两个层次上 —— 除了上面这条 command list 的关闭,提交线程在遍历 payload 时也会为它单独 flush 一次:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Submission.cpp:887
if (Payload->HasUpdateReservedResourcesWork())
{
Flush();
UpdateReservedResources(Payload);
}这条性质的连带成本见 §8。
3.4 队列能力兜底
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Submission.cpp:653
// On some devices, some queues cannot perform tile remapping operations.
// We can work around this limitation by running the remapping in lockstep on another queue:
// - tile mapping queue waits for commands on this queue to finish
// - tile mapping queue performs the commit/decommit operations
// - this queue waits for tile mapping queue to finish
ID3D12CommandQueue* TileMappingQueue = (Queue.bSupportsTileMapping ? Queue.D3DCommandQueue : Queue.Device->TileMappingQueue).GetReference();
const bool bCrossQueueSyncRequired = TileMappingQueue != Queue.D3DCommandQueue.GetReference();不支持 tile mapping 的队列(如某些硬件上的 copy / compute queue)会回退到 Direct queue(D3D12Device.cpp:135),代价是一次双向 fence 的串行化:
跨队列 tile mapping 的同步时序:两次等待缺一不可
前一次等待防止「旧访问还没做完就改页表」,后一次防止「新映射还没生效就访问新区域」。
⚠️ 这与 NVIDIA 建议的「把 UpdateTileMappings 放到 async copy queue」是相反方向的取舍 —— UE 选择跟随提交队列、必要时兜底,而不是无脑丢给 copy queue,因为 RDG 的 commit 时点与 pass 依赖强绑定,跨队列会引入额外的同步点。
四、核心实现精读:CommitReservedResource
实现在 Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp:223-611。
4.1 心智模型:一条只在尾部伸缩的水位线
这个函数只干一件事:
把「我要多少物理显存」这个字节数,翻译成一串
UpdateTileMappings调用,并顺便管理背后的 heap。
对上 C++ 容器的直觉:
CreateReservedResource(MaxSize) ≈ reserve():确定 capacity,此后不可改
CommitReservedResource(NewSize) ≈ resize() :改变当前「有物理 backing」的 size
类比到此为止 —— 资源本身永远是 reserved resource,不会被「转化」成 committed resource;而且 resize() 会保留旧数据,Commit 新增出来的区域内容是未定义的。
关键设计:UE 的 reserved resource 不是「任意稀疏」,而是「线性尾部伸缩的池」。
Reserved Resource 的 tile 空间(VA,创建时就固定,比如 32768 个 tile = 2GB)
┌─────────────────────────────────────────────────────────────────────┐
│■■■■■■■■■■■■■■■■■■■■■│░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░│
└─────────────────────┬───────────────────────────────────────────────┘
已提交(有物理内存) │ 未提交(映射到 NULL,不占显存)
NumCommittedTiles ───┘ ← 这个函数就是在左右移动这条线
没有「中间挖洞」,没有随机稀疏,只有一个标量 NumCommittedTiles。想通这一点,函数就塌掉一半难度 —— 这也解释了为什么它主要服务于「可增长的大 buffer」,而不是「virtual texture 式的随机稀疏」。
物理内存那边长这样:
NumCommittedTiles = 640
tile: [0 ────────── 255][256 ───────── 511][512 ─── 639]
└── Heap A ──┘ └── Heap B ──┘ └─ Heap C ─┘
16MB/256格 16MB/256格 8MB/128格
↑ BackingHeaps 数组,顺序 = tile 顺序
后文会把这套映射关系写成「概念页表」的形式:
PTE[R[i]] = { Heap 对象, Heap 内 tile 序号 } 或 PTE[R[i]] = NULL
这只是理解 API 语义的模型,不是 D3D12 暴露给应用的真实硬件页表结构。真实实现里 GPU 页表可能是多级的、heap 对应的 WDDM allocation 可能被分页、实际页面可能在本地 VRAM 也可能被 OS 迁移。D3D12 向应用暴露的稳定抽象只有一条:reserved resource 的 tile ↔ 某个 ID3D12Heap 里的 tile offset。
4.2 四个状态变量与六条不变量
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.h:223
struct FD3D12ReservedResourceData
{
TArray<TRefCountPtr<FD3D12Heap>> BackingHeaps;
// Flattened array of residency handles owned by backing heaps, used to support batched GetResidencyHandles()
TArray<FD3D12ResidencyHandle*> ResidencyHandles;
TArray<int32> NumResidencyHandlesPerHeap;
// Tiles currently assigned to the resource
uint32 NumCommittedTiles = 0;
// Available tiles at the end of the last backing heap
uint32 NumSlackTiles = 0;
};| 含义 | 一句话 | |
|---|---|---|
BackingHeaps | 物理内存的分段列表,顺序即 tile 顺序 | 第 i 个 heap 服务 tile 区间的第 i 段 |
ResidencyHandles | 拍平的驻留句柄,NumResidencyHandlesPerHeap 记录每个 heap 贡献了几个,方便 pop 时同步弹出 | 记账用 |
NumCommittedTiles | 水位线。tile [0, N) 有内存 | 全函数的核心状态 |
NumSlackTiles | 最后一个 heap 尾部有多少 tile 已解映射但 heap 还留着 | 见下 |
❗ NumSlackTiles 是唯一反直觉的地方。新建 heap 时,尺寸是按需精确分配的:
const uint32 ThisHeapSize = RegionSize.NumTiles * TileSizeInBytes; // 要几个 tile 就建多大所以新建的 heap 永远没有 slack。NumSlackTiles > 0 只可能由「收缩」产生:收缩时如果某个 heap 只被解映射了一部分,这个 heap 不能删(前半段还在用),于是它尾部就空出来一段 —— 记成 slack,下次增长优先复用,避免「删了又建」的抖动。
收缩前: [Heap B: ████████████████] 256 tiles 全用
收缩后: [Heap B: █████░░░░░░░░░░░] 用了 44,slack = 212(heap 还在,内存还占着)
再增长: [Heap B: █████████░░░░░░░] 从 offset 44 复用 100 个,slack = 112
必须维持的六条不变量
整个函数的读法可以简化成「检查这六条永远成立」:
- 已映射区域始终是
[0, NumCommittedTiles); - 资源中间不留洞;
BackingHeaps按资源虚拟 tile 顺序排列(第 i 个 heap 服务第 i 段);- 只有最后一个 heap 可以含 slack;
- 扩容和缩容都只从尾部进行;
- 保留下来的前缀仍指向原来的 heap tile,因此永远不需要复制数据。
第 3、4 条是 §4.4 那两段 while 循环能成立的全部前提 —— 正因为 slack 只可能出现在最后一个 heap,HeapFirstTile = NumCommittedTiles - NumUsedTilesInHeap 这种「从水位线往回倒推 heap 起点」的算法才是对的。
由此得到两个记账公式:
CommittedBytes = NumCommittedTiles × 64KB
PhysicalHeapBytes = CommittedBytes + NumSlackTiles × 64KB
第二个公式解释了一个实际现象:stat rhi 里的 STAT_D3D12ReservedResourcePhysical(backing heap 实际字节数)可能大于逻辑上已提交的字节数,差额就是 slack。这个统计量在 CreateHeap 时 INC、在 DeferDelete() 时 DEC,跟的是 heap 的生死,而不是水位线。
4.3 函数骨架:六个阶段
void FD3D12Resource::CommitReservedResource(ID3D12CommandQueue* D3DCommandQueue, uint64 RequiredCommitSizeInBytes)
{
// 阶段 1:前置检查
// 阶段 2:问 D3D「这个资源是怎么切 tile 的」
// 阶段 3:算账(目标 tile 数、heap 尺寸上限)
// 阶段 4:定义「线性 tile 序号 → D3D 坐标」的换算 lambda
// 阶段 5:走 收缩 或 增长 分支,产出 MappingParams 数组(此时还没调 D3D)
// 阶段 6:MakeResident → 批量 UpdateTileMappings → SignalFence → 更新统计
}画成三段结构(阶段 4 只是定义一个 lambda,不占流程上的位置,故未画出):
阶段 5 只往数组里记条目,一次 D3D 都不调;真正的执行集中在阶段 6
5A 和 5B 会按 heap 逐轮循环(一次收缩或增长可能跨越多个 heap),两条分支最终汇合到同一处,MakeResident → UpdateTileMappings → SignalFence 只在阶段 6 发生一次。
之所以只「记账」不「执行」:MakeResident 必须发生在 UpdateTileMappings 之前(不能把 tile 映射到一个还没驻留的 heap),而新 heap 是在阶段 5 里才创建的。所以必须先全部算完、收集齐 residency handle,再统一执行。
4.4 逐阶段精读
阶段 1:前置检查(L225-242)
TRACE_CPUPROFILER_EVENT_SCOPE(CommitReservedResource);
static constexpr uint64 TileSizeInBytes = GRHIGlobals.ReservedResources.TileSizeInBytes;
static_assert(TileSizeInBytes == 65536, "Reserved resource tiles are expected to always be 64KB");
check(Desc.bReservedResource);
check(ReservedResourceData.IsValid());
LLM_REALLOC_SCOPE(ReservedResourceData.Get()); // 内存统计:把后面 CreateHeap 的分配挂到这个资源名下
UE_MEMSCOPE_PTR(ReservedResourceData.Get());
checkf(GRHIGlobals.ReservedResources.Supported, ...);
if (Desc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE3D)
checkf(GRHIGlobals.ReservedResources.SupportsVolumeTextures, ...); // 3D 需要 Tier 3LLM_REALLOC_SCOPE(ReservedResourceData.Get()) 用这个结构体的地址当作 LLM 的分配 key —— 因为 reserved 资源的「地址」会随着 heap 增删而变,没法用 GPU VA 当 key。
顺带一提,函数开头那个 TRACE_CPUPROFILER_EVENT_SCOPE(CommitReservedResource) 就是后面 §9 里在 Insights 判定「是否真的走了 reserved 路径」的抓手。
阶段 2:问 D3D 资源怎么切 tile(L244-272)
uint32 D3DResourceNumTiles = 0; // out: 整个资源一共多少 tile
D3D12_PACKED_MIP_INFO PackedMipDesc = {}; // out: 标准 mip / 打包 mip 的分界
D3D12_TILE_SHAPE TileShape = {}; // out: 一个 tile 是多少 texel(这里查了但没用)
uint32 NumSubresourceTilings = NumMipLevels; // in/out
TArray<D3D12_SUBRESOURCE_TILING, TInlineAllocator<16>> MipTilingInfo;
MipTilingInfo.SetNum(NumSubresourceTilings);
D3DDevice->GetResourceTiling(GetResource(), &D3DResourceNumTiles, &PackedMipDesc,
&TileShape, &NumSubresourceTilings, /*FirstSubresource*/0,
MipTilingInfo.GetData());两个返回结构:
struct D3D12_PACKED_MIP_INFO {
UINT8 NumStandardMips; // 前 N 级 mip 是「标准 tile 排布」
UINT8 NumPackedMips; // 后面这些小 mip 被打包进 mip tail
UINT NumTilesForPackedMips; // mip tail 占几个 tile(每个 array slice)
UINT StartTileIndexInOverallResource;
};
struct D3D12_SUBRESOURCE_TILING { // 每级标准 mip 一个
UINT WidthInTiles, StartTileIndexInOverallResource;
UINT16 HeightInTiles, DepthInTiles;
};这里体现了 §2.1 那条「不要硬编码 tile 形状」的原则:一切以运行时查询为准。另外注意它只查了 array slice 0 的 mip 链(FirstSubresource = 0,数量 = NumMipLevels)—— 因为所有 slice 的 tiling 完全一样,后面用乘法补回 slice 维度。
然后是那个「作弊」:
if (bBuffer)
{
// Buffers obviously don't have mips, but we can pretend they do to make the code below agnostic to resource type
PackedMipDesc.NumStandardMips = 1;
}
check(MipTilingInfo.Num() == PackedMipDesc.NumStandardMips + PackedMipDesc.NumPackedMips);buffer 没有 mip,D3D 返回的 PackedMipDesc 全 0。UE 强行填 NumStandardMips = 1,把 buffer 伪装成「1 级标准 mip 的纹理」,这样下面所有代码只有一套逻辑,不需要 if/else 分叉。
阶段 3 里还有一条与之配套的断言,把这个伪装的前提钉死:
checkf(D3DResourceNumTiles == MipTilingInfo[0].WidthInTiles,
TEXT("Reserved buffers are expected to have trivial tiling configuration: single 1D subresource that contains all tiles."));阶段 3:算账(L274-324)
const uint32 NumPackedTilesPerArraySlice = PackedMipDesc.NumTilesForPackedMips;
const uint32 NumTotalPackedMipTiles = NumPackedTilesPerArraySlice * NumArraySlices;
const uint32 NumTotalStandardMipTiles = D3DResourceNumTiles - NumTotalPackedMipTiles;
const uint64 TotalSize = D3DResourceNumTiles * TileSizeInBytes;
RequiredCommitSizeInBytes = FMath::Min<uint64>(RequiredCommitSizeInBytes, TotalSize); // 传 UINT64_MAX = 全提交
RequiredCommitSizeInBytes = AlignArbitrary(RequiredCommitSizeInBytes, TileSizeInBytes); // 上取整到 64KB
const uint64 MaxHeapSize = uint64(CVarD3D12ReservedResourceHeapSizeMB.GetValueOnAnyThread()) * 1024 * 1024; // 默认 16MB
const uint64 NumHeaps = FMath::DivideAndRoundUp(TotalSize, MaxHeapSize); // 仅用于 BackingHeaps.Reserve()
const uint32 MaxTilesPerHeap = uint32(MaxHeapSize / TileSizeInBytes); // 默认 256
const uint32 NumRequiredCommitTiles = RequiredCommitSizeInBytes / TileSizeInBytes; // ★ 目标水位线UINT64_MAX 是「全部提交」的约定值(TexCreate_ImmediateCommit 就走这条)。
物理内存不是一个大 heap,而是一串默认 16MB 的小 heap:
// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp:16
static TAutoConsoleVariable<int32> CVarD3D12ReservedResourceHeapSizeMB(
TEXT("d3d12.ReservedResourceHeapSizeMB"), 16,
TEXT("Size of the backing heaps for reserved resources in megabytes (default 16MB)."), ECVF_ReadOnly);Vulkan 后端有一个一模一样的旋钮 r.Vulkan.SparseImageAllocSizeMB(VulkanTexture.cpp:43),默认值同样是 16MB,同样 ECVF_ReadOnly。
heap 的 residency 优先级沿用 committed resource 的启发式,高优先级的会额外调一次 SetResidencyPriority(D3D12_RESIDENCY_PRIORITY_HIGH):
const bool bRenderOrDepthTarget = Flags & (ALLOW_RENDER_TARGET | ALLOW_DEPTH_STENCIL);
const bool bHighPriorityResource = bRenderOrDepthTarget || (Flags & ALLOW_UNORDERED_ACCESS);heap flag 也按用途收紧(这会影响某些硬件的压缩 / 对齐行为):
const D3D12_HEAP_FLAGS TextureHeapFlags = bRenderOrDepthTarget
? D3D12_HEAP_FLAG_ALLOW_ONLY_RT_DS_TEXTURES
: D3D12_HEAP_FLAG_ALLOW_ONLY_NON_RT_DS_TEXTURES;
const D3D12_HEAP_FLAGS HeapFlags = bBuffer ? D3D12_HEAP_FLAG_ALLOW_ONLY_BUFFERS : TextureHeapFlags;阶段 5A:收缩分支(L395-453)
if (ReservedResourceData->NumCommittedTiles > NumRequiredCommitTiles) // 要变小
{
while (NumCommittedTiles > NumRequiredCommitTiles)
{
// ── 1. 定位「最后一个 heap」覆盖的 tile 区间 ──
TRefCountPtr<FD3D12Heap>& LastHeap = BackingHeaps.Last();
const uint32 NumTotalTilesInHeap = LastHeap->GetHeapDesc().SizeInBytes / TileSizeInBytes;
const uint32 NumUsedTilesInHeap = NumTotalTilesInHeap - NumSlackTiles;
const uint32 HeapFirstTile = NumCommittedTiles - NumUsedTilesInHeap; // 这个 heap 从哪个资源 tile 开始
// ── 2. 本轮要解映射的区间 = [max(HeapFirstTile, 目标), 当前水位) ──
const uint32 RegionEnd = NumCommittedTiles;
const uint32 RegionBegin = FMath::Max(HeapFirstTile, NumRequiredCommitTiles);
D3D12_TILE_REGION_SIZE RegionSize = {};
RegionSize.UseBox = false; // 线性模式,可跨 mip/slice
RegionSize.NumTiles = RegionEnd - RegionBegin;
// ── 3. 记一条「映射到 NULL」的指令 ──
FD3D12UpdateTileMappingsParams Params = {};
Params.RangeFlags = D3D12_TILE_RANGE_FLAG_NULL; // Heap 保持 nullptr
Params.Coord = GetTiledResourceCoordinate(RegionBegin, RegionSize.NumTiles);
Params.Size = RegionSize;
MappingParams.Add(Params);
// ── 4. 这个 heap 是全空了还是半空 ──
if (HeapFirstTile == RegionBegin) // 整个 heap 都被解掉 → 释放
{
DEC_MEMORY_STAT_BY(STAT_D3D12ReservedResourcePhysical, LastHeap->GetHeapDesc().SizeInBytes);
LastHeap->DeferDelete();
BackingHeaps.Pop();
/* 同步弹出对应数量的 residency handle */
NumSlackTiles = 0; // 新的 Last 是满的
}
else // 只解了尾巴 → 留着,记 slack
{
NumSlackTiles += RegionSize.NumTiles;
}
NumCommittedTiles -= RegionSize.NumTiles;
}
}为什么是 while 循环? 因为一次收缩可能跨越多个 heap(比如从 640 掉到 300,要先整个干掉 Heap C,再切掉 Heap B 的一部分)。每轮只处理「当前最后一个 heap」。
HeapFirstTile 那行是全函数最容易看晕的:
NumCommittedTiles = 640, LastHeap = C(128 tiles), NumSlackTiles = 0
NumUsedTilesInHeap = 128 - 0 = 128
HeapFirstTile = 640 - 128 = 512 ← Heap C 服务的是资源 tile [512, 640)
本质就是「从水位线往回倒推最后一个 heap 的起点」。
阶段 5B:增长分支(L454-551)
else // 要变大(或不变,循环直接不进)
{
while (NumCommittedTiles < NumRequiredCommitTiles)
{
const uint32 NumRemainingTiles = NumRequiredCommitTiles - NumCommittedTiles;
ID3D12Heap* D3DHeap = nullptr;
uint32 HeapRangeStartOffsetInTiles = 0;
D3D12_TILE_REGION_SIZE RegionSize = {};
RegionSize.UseBox = false;
if (NumSlackTiles) // ── 路径 A:先吃掉现有 heap 的余量 ──
{
const auto& LastHeap = BackingHeaps.Last();
const uint32 NumTotalTilesInHeap = LastHeap->GetHeapDesc().SizeInBytes / TileSizeInBytes;
RegionSize.NumTiles = FMath::Min(NumSlackTiles, NumRemainingTiles);
HeapRangeStartOffsetInTiles = NumTotalTilesInHeap - NumSlackTiles; // ★ 从 heap 里第几个 tile 开始
D3DHeap = LastHeap->GetHeap();
NumSlackTiles -= RegionSize.NumTiles;
UsedResidencyHandles.Append(LastHeap->GetResidencyHandles());
}
else // ── 路径 B:新建一个 heap ──
{
RegionSize.NumTiles = FMath::Min(MaxTilesPerHeap, NumRemainingTiles); // 一次最多 16MB
HeapRangeStartOffsetInTiles = 0;
NewHeapDesc.SizeInBytes = RegionSize.NumTiles * TileSizeInBytes; // 精确尺寸,无 slack
VERIFYD3D12RESULT(D3DDevice->CreateHeap(&NewHeapDesc, IID_PPV_ARGS(&D3DHeap)));
INC_MEMORY_STAT_BY(STAT_D3D12ReservedResourcePhysical, NewHeapDesc.SizeInBytes);
/* 包装成 FD3D12Heap、BeginTrackingResidency、追加到 BackingHeaps 和 handle 表 */
}
FD3D12UpdateTileMappingsParams Params = {};
Params.RangeFlags = D3D12_TILE_RANGE_FLAG_NONE; // 顺序映射
Params.Coord = GetTiledResourceCoordinate(NumCommittedTiles, RegionSize.NumTiles);
Params.Size = RegionSize;
Params.Heap = D3DHeap;
Params.HeapOffsetInTiles = HeapRangeStartOffsetInTiles;
MappingParams.Add(Params);
NumCommittedTiles += RegionSize.NumTiles;
}
}HeapRangeStartOffsetInTiles = NumTotalTilesInHeap - NumSlackTiles 为什么对?§4.6 会用具体数字验算。
阶段 6:真正调 D3D(L553-610)
// ① 先让所有用到的 heap 驻留
if (GEnableResidencyManagement && !UsedResidencyHandles.IsEmpty())
{
FD3D12ResidencySet* ResidencySet = ResidencyManager.CreateResidencySet();
ResidencySet->Open();
for (FD3D12ResidencyHandle* Handle : UsedResidencyHandles) ResidencySet->Insert(Handle);
ResidencySet->Close();
ResidencyManager.MakeResident(D3DCommandQueue, MoveTemp(ResidencySet));
}
// ② 再逐条改页表
for (const FD3D12UpdateTileMappingsParams& Params : MappingParams)
{
D3DCommandQueue->UpdateTileMappings(GetResource(),
1 /*NumRegions*/, &Params.Coord, &Params.Size,
Params.Heap,
1 /*NumRanges*/, &Params.RangeFlags, &Params.HeapOffsetInTiles, &Params.Size.NumTiles,
D3D12_TILE_MAPPING_FLAG_NONE);
}
// ③ 告诉 residency manager「这批 heap 已被队列引用」
// Signal the fence for this queue after UpdateTileMappings complete.
// This is analogous to executing a command list that references a set of resources.
ResidencyManager.SignalFence(D3DCommandQueue);三步的顺序不能换:因为 reserved resource 本身不做 residency 追踪(见 §3.2 的 NOTE),驻留单位是 backing heap,必须先 MakeResident 才能映射过去。
这里必须分清两个不同的问题(对应 §2.1 分层图的 tile mapping 层与 residency 层):
| 回答的问题 | 由谁负责 | |
|---|---|---|
| Tile Mapping | Resource virtual tile → 哪个 Heap 的哪个 tile? | ID3D12CommandQueue::UpdateTileMappings() |
| Residency | 提供 backing 的这个 Heap,当前在不在 GPU 可访问的驻留集合里? | UE 的 Residency Manager |
正确时序只有一条:创建 Heap → 让 Heap resident → 建立 tile 映射 → GPU 才能安全访问新区域。最后那次 SignalFence 的作用,是告诉 Residency Manager「这批 heap 已被这个 queue 引用到了这个时间点」,避免它过早把 heap 移出驻留集合或回收。
这个 10 参数的签名,按「左右两侧」读就清楚了:
| 参数 | 含义 | |
|---|---|---|
| 左侧:Resource 虚拟 tile 区间 | pResource / pResourceRegionStartCoordinates / pResourceRegionSizes | 从资源的 Coord 开始,连续取 N 个虚拟 tile |
| 右侧:Heap tile 区间 | pHeap / pRangeFlags / pHeapRangeStartOffsets / pRangeTileCounts | 从指定 Heap 的 HeapOffset 开始,连续取 N 个 backing tile |
一次普通映射的精确语义,就是把两侧一一对应起来:
Map(ResourceStartTile = X, NumTiles = N, Heap = HeapK, HeapStartTile = H)
对 k = 0 … N-1: R[X + k] → HeapK[H + k]
而 RangeFlags = NULL 时右侧整个不存在,语义退化成 R[X + k] → NULL。它改的自始至终只是地址翻译关系,不搬运任何数据 —— 这也是「新提交区域内容未定义」的根本原因。
参数配对关系上,「区域侧的总 tile 数」必须等于「范围侧的总 tile 数」。UE 每次只发一个区域 + 一个范围,所以直接把同一个 Params.Size.NumTiles 同时当作区域大小和范围计数传进去 —— 这就是 &Params.Size.NumTiles 出现在 pRangeTileCounts 位置的原因。
⚠️ 这里其实可以把 MappingParams 打包成一次 UpdateTileMappings(NumRegions=N, NumRanges=N) 调用,减少驱动调用次数。UE 选了逐条发,应该是为了代码简单 + 单次 commit 的 params 通常只有 1~3 条(16MB 粒度已经把次数压得很低了)。
最后更新统计:
if (ReservedResourceData->NumCommittedTiles != NumPreviousCommittedTiles)
{
int64 CommitDeltaInBytes = TileSizeInBytes * FMath::Abs((int32)NumCommittedTiles - (int32)NumPreviousCommittedTiles);
UE::RHICore::UpdateReservedResourceStatsOnCommit(CommitDeltaInBytes, bBuffer,
NumCommittedTiles > NumPreviousCommittedTiles);
}UpdateReservedResourceStatsOnCommit 干的事是把这批字节数从 ReservedUncommitted* 挪到 ReservedCommitted*(RHICoreStats.cpp:195)—— 一个 INC 一个 DEC,所以两个统计量之和恒等于预留总量。§9 就是靠这个差值判断省了多少显存。
4.5 GetTiledResourceCoordinate:线性序号 → D3D 坐标
这个 lambda 回答一个问题:「资源里第 K 个 tile」在 D3D 眼里是 (Subresource, X, Y, Z) 的哪一个?
D3D 不认识「第 K 个 tile」,它只认识「第几个 subresource 的第 (x,y,z) 个 tile」。而 UE 的整套水位线逻辑是线性的,所以必须有这个翻译层。
先破除一个由函数名引起的误解:它只计算 Resource 侧的坐标,既不返回 Heap,也不返回 heap offset。 映射的两侧是由三个不同参数分别描述的:
| 参数 | 属于哪个坐标空间 |
|---|---|
D3D12_TILED_RESOURCE_COORDINATE Coord | Resource 虚拟 tile 坐标 —— 这个 lambda 的唯一产物 |
HeapOffsetInTiles | Heap 内的 backing tile 起点 —— 由阶段 5 的两条路径决定 |
NumTiles | 两侧连续对应的数量 |
前提:tile 的线性排布规则
资源整体(假设 2 个 array slice)
┌──────────────────── slice 0 ────────────────────┬──────── slice 1 ────────┐
│ mip0 tiles │ mip1 tiles │ mip2 tiles │ mip tail │ mip0 │ mip1 │ ... │tail │
└─────────── NumTotalTilesPerArraySlice ──────────┴─────────────────────────┘
所以:
const uint32 ArraySliceIndex = OffsetInTiles / NumTotalTilesPerArraySlice;
const uint32 TileIndexInArraySlice = OffsetInTiles % NumTotalTilesPerArraySlice;第一步:确定落在哪一级 mip
uint32 MipLevel = 0;
uint32 NextMipTileThreshold = 0;
while (MipLevel < PackedMipDesc.NumStandardMips)
{
const D3D12_SUBRESOURCE_TILING& CurrentMipTiling = MipTilingInfo[MipLevel];
NextMipTileThreshold += CurrentMipTiling.WidthInTiles * CurrentMipTiling.HeightInTiles * CurrentMipTiling.DepthInTiles;
if (TileIndexInArraySlice < NextMipTileThreshold) break; // 落在本级
MipLevel += 1;
}如果一路没 break,循环自然停在 MipLevel == NumStandardMips —— 这正好就是「第一个 packed mip」的索引,下面靠这个区分两条路。
ResourceCoordinate.Subresource = MipLevel + ArraySliceIndex * NumTotalMips;D3D12 的 subresource 编号规则就是 MipSlice + ArraySlice * MipLevels。
第二步 · 情况 A:标准 mip
const uint32 NumTilesPerVolumeSlice = CurrentMipTiling.WidthInTiles * CurrentMipTiling.HeightInTiles;
const uint32 TileIndexInMipLevel = TileIndexInArraySlice - CurrentMipTiling.StartTileIndexInOverallResource;
ResourceCoordinate.X = TileIndexInMipLevel % CurrentMipTiling.WidthInTiles;
ResourceCoordinate.Y = (TileIndexInMipLevel / CurrentMipTiling.WidthInTiles) % CurrentMipTiling.HeightInTiles;
ResourceCoordinate.Z = TileIndexInMipLevel / NumTilesPerVolumeSlice;就是标准的一维序号 → 三维下标解包(行主序)。StartTileIndexInOverallResource 是这级 mip 在 slice 内的起始 tile 号,减掉它就得到 mip 内的局部序号。
第二步 · 情况 B:packed mip(mip tail)
checkf(NumTiles <= MaxTilesPerHeap,
TEXT("Reserved texture packed mip level requires tiles: %d, maximum supported tiles: %d. ")
TEXT("Increase d3d12.ReservedResourceHeapSizeMB or avoid packed mips by using a larger texture dimensions."),
NumTiles, MaxTilesPerHeap);
// Entire packed mip chain must be covered in one map operation, so mapping origin is always 0
ResourceCoordinate.X = 0;
ResourceCoordinate.Y = 0;
ResourceCoordinate.Z = 0;mip tail 没有几何形状可言(驱动内部怎么塞的是黑箱),只能整体映射。官方规定是:当区域包含非标准 tiling 的 mip 时 UseBox 必须为 FALSE,起始位置用 x 在扁平 tile 范围里偏移、y 和 z 必须为 0。UE 因为总是整条一起映射,所以 x 也固定为 0。
那句 assert 的实际含义:mip tail 必须能被一个 heap 一次性覆盖(因为一次 UpdateTileMappings 只能指向一个 heap)。这正是 §2.2 那条「packed mip 必须一次映射完」的直接后果。如果 mip tail 超过 16MB,就会撞上这个 check —— 报错信息也直接给了两条出路:调大 d3d12.ReservedResourceHeapSizeMB,或者把纹理做大到不产生 packed mip。
顺带解释这个 lambda 的第二个参数 NumTiles:它在 buffer 与标准 mip 路径下完全不影响返回的坐标,唯一的用处就是上面这句 packed-mip 的约束检查。第一次读的时候可以当它不存在。
4.6 数字走查(一):Buffer 的增长 → 收缩 → 再增长
设定:2GB reserved buffer,d3d12.ReservedResourceHeapSizeMB = 16
TileSizeInBytes = 65536 (64KB)
D3DResourceNumTiles = 2GB / 64KB = 32768
MaxTilesPerHeap = 16MB / 64KB = 256
bBuffer = true → NumStandardMips 强制为 1
MipTilingInfo[0] = { WidthInTiles=32768, Height=1, Depth=1, Start=0 }
NumTotalTilesPerArraySlice = 32768
① 首次提交 40MB(0 → 640 tiles)
NumRequiredCommitTiles = 40MB / 64KB = 640
| 轮次 | slack | 动作 | RegionSize | Coord (X) | Heap / offset | 之后 NumCommittedTiles |
|---|---|---|---|---|---|---|
| 1 | 0 | 建 Heap A (16MB) | 256 | GTRC(0) → X=0 | A / 0 | 256 |
| 2 | 0 | 建 Heap B (16MB) | 256 | GTRC(256) → X=256 | B / 0 | 512 |
| 3 | 0 | 建 Heap C (8MB) | 128 | GTRC(512) → X=512 | C / 0 | 640 |
第三个 heap 只建了 8MB(128 × 64KB),不是 16MB —— 所以此时 slack 仍为 0。
坐标怎么算的(以 GTRC(256) 为例):ArraySliceIndex = 256/32768 = 0;mip 循环第一轮 NextMipTileThreshold = 32768,256 < 32768 成立直接 break → MipLevel = 0;Subresource = 0;X = (256 - 0) % 32768 = 256,Y = 0,Z = 0。
此时的概念页表:
R[0] → Heap A[0] R[256] → Heap B[0] R[512] → Heap C[0]
R[1] → Heap A[1] R[257] → Heap B[1] R[513] → Heap C[1]
... ... ...
R[255] → Heap A[255] R[511] → Heap B[255] R[639] → Heap C[127]
R[640] … R[32767] → NULL / unmapped
NumCommittedTiles = 640, NumSlackTiles = 0
①′ 反过来看:一次读访问是怎么翻译过去的
上面全是「怎么建立映射」。反过来,shader 读这个 buffer 的第 17MB 处会发生什么:
ByteOffset = 17MB
TileIndex = 17MB / 64KB = 272
页表查询: R[272] → Heap B[272 - 256] = Heap B[16]
Heap 内偏移:16 × 64KB = 1MB
完整翻译路径:
ReservedBuffer GPUVA + 17MB
↓
Resource virtual tile R[272]
↓ tile mapping
Heap B 的第 16 个 tile
↓
Heap B 内偏移 1MB
Shader 只看到一个连续的 buffer,完全不知道 17MB 处已经跨到了第二个 heap。 这就是 reserved resource 全部收益的来源:物理侧可以碎成 N 段、可以随时增删,虚拟侧的地址和 descriptor 岿然不动。
② 收缩到 300 tiles(640 → 300)
| 轮次 | LastHeap | NumTotalTilesInHeap | slack(前) | Used | HeapFirstTile | RegionBegin | 解映射 tile 数 | heap 处置 | slack(后) | 水位 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | C | 128 | 0 | 128 | 640-128=512 | max(512,300)=512 | 640-512=128 | 512 == 512 → 删除 C | 0 | 512 |
| 2 | B | 256 | 0 | 256 | 512-256=256 | max(256,300)=300 | 512-300=212 | 256 != 300 → 保留 B | 212 | 300 |
结果:BackingHeaps = [A, B],NumCommittedTiles = 300,NumSlackTiles = 212。Heap B 只有前 44 个 tile 在用(对应资源 tile 256~299)。
❗ 这里必须分清两件事:
R[300..511] → NULL ✅ 发生了:资源侧的映射被清除
Heap B[44..255] 被删除 ❌ 没有发生:这段物理内存仍属于 Heap B,只是没有资源 tile 指向它了
D3D12_TILE_RANGE_FLAG_NULL 清除的是资源 → heap 的连接,不销毁 heap 对象。只有当整个尾部 heap 都不再被任何资源 tile 指向时(本例中的 Heap C),UE 才 DeferDelete() 它。这也正是 NumSlackTiles = 212 的物理含义:已经付了钱、但暂时没人用的 13.25MB。
③ 再增长到 400 tiles(300 → 400)
NumRemainingTiles = 100,NumSlackTiles = 212 > 0 → 走「路径 A:吃 slack」
RegionSize.NumTiles = min(212, 100) = 100
HeapRangeStartOffsetInTiles = 256 - 212 = 44
Coord = GTRC(300) → X = 300
NumSlackTiles = 212 - 100 = 112
NumCommittedTiles = 300 + 100 = 400
验算这个 44 对不对:Heap B 服务的是资源 tile [256, 512)。资源 tile 300 应该落在 Heap B 的第 300 - 256 = 44 个 tile。完全吻合,一次 CreateHeap 都没多花。
4.7 数字走查(二):纹理的坐标换算
设定:512×512、RGBA8(32bpp)、10 级完整 mip、非数组。32bpp 的标准 tile 是 128×128 texel。
mip0 512×512 → 4×4 = 16 tiles, Start = 0
mip1 256×256 → 2×2 = 4 tiles, Start = 16
mip2 128×128 → 1×1 = 1 tile, Start = 20
mip3~9 (64×64 及以下) → 小于一个 tile → 打包成 mip tail,占 1 tile
NumStandardMips = 3, NumPackedMips = 7, NumTilesForPackedMips = 1
D3DResourceNumTiles = 16 + 4 + 1 + 1 = 22
NumTotalTilesPerArraySlice = 21 + 1 = 22
NumTotalMips = MipTilingInfo.Num() = 3 + 7 = 10
GetTiledResourceCoordinate(18)
ArraySliceIndex = 18 / 22 = 0
TileIndexInArraySlice = 18 % 22 = 18
mip 循环:
MipLevel=0: threshold += 16 → 16; 18 < 16 ? 否 → MipLevel=1
MipLevel=1: threshold += 4 → 20; 18 < 20 ? 是 → break
→ MipLevel = 1,Subresource = 1 + 0×10 = 1
标准 mip 路径(1 < 3):
T = mip1 { Width=2, Height=2, Depth=1, Start=16 }
TileIndexInMipLevel = 18 - 16 = 2
X = 2 % 2 = 0
Y = (2 / 2) % 2 = 1
Z = 2 / (2×2) = 0
→ (Subresource=1, X=0, Y=1, Z=0)
翻译成人话:mip1 里第 2 行第 1 列的那个 128×128 区块。mip1 是 256×256,正好 2×2 个 tile,index 2 就是左下角。
GetTiledResourceCoordinate(21)
TileIndexInArraySlice = 21
mip 循环:
MipLevel=0: threshold=16; 21<16 否 → 1
MipLevel=1: threshold=20; 21<20 否 → 2
MipLevel=2: threshold=21; 21<21 否 → 3
MipLevel(3) < NumStandardMips(3) ? 否 → 循环退出
→ MipLevel = 3 = NumStandardMips → 走 packed 分支
→ (Subresource=3, X=0, Y=0, Z=0),并断言这次 NumTiles ≤ MaxTilesPerHeap
Subresource 3 就是第一个 packed mip,坐标 (0,0,0) + 线性 NumTiles 覆盖整条 mip tail。
五、Buffer 与 Texture 的区分
上一节的水位线模型对 buffer 完美适用。但对纹理,UE 只支持创建时一次性全量提交(TexCreate_ImmediateCommit),运行时不能 commit / decommit。这不是文档里的一句话,是四层代码同时堵死的。
5.1 四条硬证据
证据 1:RHI 验证层直接把纹理堵死了
// Engine/Source/Runtime/RHI/Private/RHIValidation.cpp:930-944
if (const FRHICommitResourceInfo* CommitInfo = Info.CommitInfo.GetPtrOrNull())
{
if (Info.Type == FRHITransitionInfo::EType::Buffer)
{
RHI_VALIDATION_CHECK(EnumHasAllFlags(BufferUsage, BUF_ReservedResource), ...);
RHI_VALIDATION_CHECK(CommitInfo->SizeInBytes <= BufferSize, ...);
}
else
{
RHI_VALIDATION_CHECK(false, TEXT("Reserved resource commit is only supported for buffers"));
}
}证据 2:D3D12 后端连分支都没写
// Engine/Source/Runtime/D3D12RHI/Private/D3D12LegacyBarriers.cpp:1185
// Enhanced Barriers 版本在 D3D12EnhancedBarriers.cpp:3303 是逐字相同的一份
if (Info.Type == FRHITransitionInfo::EType::Buffer)
{
FD3D12Buffer* Buffer = Context.RetrieveObject<FD3D12Buffer>(Info.Buffer);
Context.SetReservedBufferCommitSize(Buffer, CommitInfo->SizeInBytes);
}
else
{
checkNoEntry(); // ← 纹理走到这里直接崩
}而且函数名就叫 SetReserved**Buffer**CommitSize(FD3D12**Buffer*** Buffer, ...),签名层面就只收 buffer。FRHITransitionInfo 里带 FRHICommitResourceInfo 的构造函数也只有 buffer 版本(RHITransition.h:178)。
证据 3:Vulkan 后端在实现里直接断言
// Engine/Source/Runtime/VulkanRHI/Private/VulkanTexture.cpp:931
TArray<VkSparseMemoryBind> FVulkanTexture::CommitReservedResource(uint64 RequiredCommitSizeInBytes)
{
checkf((RequiredCommitSizeInBytes == UINT64_MAX) || (RequiredCommitSizeInBytes == 0),
TEXT("Only full resource commits are supported for images."));证据 4:Vulkan 创建时干脆没开 residency bit
// Engine/Source/Runtime/VulkanRHI/Private/VulkanTexture.cpp:382
if (EnumHasAnyFlags(UEFlags, TexCreate_ReservedResource))
{
ImageCreateInfo.flags |= VK_IMAGE_CREATE_SPARSE_BINDING_BIT;
// no VK_IMAGE_CREATE_SPARSE_RESIDENCY_BIT yet since only TexCreate_ImmediateCommit is supported
}Vulkan 里这两个 bit 分工很清楚:SPARSE_BINDING_BIT = 内存可以来自多块分配,但整张图必须全部绑定;SPARSE_RESIDENCY_BIT = 允许部分驻留(真稀疏)。UE 只开了前者。
旁证:所有纹理用例无一例外都带 ImmediateCommit
| 用例 | 创建 flag |
|---|---|
VSM Physical Page Pool(VirtualShadowMapCacheManager.cpp:1096) | TexCreate_ReservedResource | TexCreate_ImmediateCommit |
SVT Tile Data Texture(SparseVolumeTextureTileDataTexture.cpp:204) | TexCreate_ReservedResource | TexCreate_ImmediateCommit |
| MediaTextureResource | 只是个透传参数,注释里提了一嘴 |
5.2 根因 A:改分辨率会打乱整个 tile↔像素映射
假设真能把 512×512「扩」成 1024×1024:
| 512×512 | 1024×1024 | |
|---|---|---|
| mip0 tile 排布 | 4×4 = 16 | 8×8 = 64 |
| tile #4 代表什么 | mip0 的 (0,1) ← 第 2 行第 1 列 | mip0 的 (4,0) ← 第 1 行第 5 列 |
| mip1 起点 | tile 16 | tile 64 |
已有数据全部错位。而 buffer 完全不同:
buffer: 元素 i 的地址 = base + i × stride ← 扩容不改变任何已有元素的位置
texture: tile k 的含义 = f(k, 分辨率, mip 链) ← 改分辨率 = 换了一个函数
「尾部追加」这个操作在纹理上没有对应的语义。
5.3 根因 B(更致命):尾部伸缩的方向和 mip streaming 是反的
就算不改分辨率,回看 tile 在资源里的排列顺序:
tile: 0 ─────────────── 15 │ 16 ─ 19 │ 20 │ 21
└──── mip0 (16) ────┘ └mip1(4)┘ └m2┘ └tail┘
↑ 占 73% 空间,最该被卸载 最该常驻 ↑
│
UE 的水位线只能从这头砍 ──┘
mip0 在最前面,mip tail 在最后面。
mip streaming 想省内存,要卸载的是 mip0(占了 73% 的 tile,且远处物体根本用不到);要保留的是小 mip(永远得有个兜底)。而 UE 的模型是 NumCommittedTiles 尾部伸缩 —— 能砍的恰恰是最该留的那些,该砍的在最前面动不了。
要支持 mip streaming,就得把标量水位线换成真正的稀疏位图 + 空闲 tile 分配器,还得处理 Tier 语义(读 NULL tile 返回 0)、shader 侧的 LOD clamp / CheckAccessFullyMapped 反馈 —— 那是另一个量级的系统。UE 没有为此付这个成本。
5.4 那 CommitReservedResource 里的 mip 逻辑给谁用?
给「分段全量提交」用的,不是给「部分驻留」用的。
关键在于 heap 上限只有 16MB(d3d12.ReservedResourceHeapSizeMB)。一个 VSM page pool 动辄 256MB512MB,全量提交也得**拆成 1632 次 UpdateTileMappings**,每次的起点是不同的线性 tile 序号 —— 而 D3D 只认 (Subresource, X, Y, Z)。
所以 GetTiledResourceCoordinate 的真实职责是:
「第 4096 个 tile 在哪级 mip 的哪个位置?」—— 为了让第 17 次
UpdateTileMappings知道从哪儿接着铺。
用 §4.7 的 512×512 例子(22 个 tile)走一遍 ImmediateCommit(RequiredCommitSizeInBytes = UINT64_MAX):
NumRequiredCommitTiles = 22(全部)
MaxTilesPerHeap = 256 → 22 < 256,一次就够
→ 建 1 个 heap(22 × 64KB = 1408KB)
→ 1 次 UpdateTileMappings:Coord = GTRC(0) = (Subresource=0, 0,0,0),NumTiles = 22
→ UseBox=false,运行时自动从 mip0 一路铺到 mip tail
只有当资源大到超过 16MB(>256 tiles)才会分多次,那时才真正用上坐标换算。VSM 的 256MB page pool 就是这种情况。
所以那套 mip 逻辑不是「稀疏能力的残留」,是「分段提交的必需品」。
5.5 纹理用 reserved 图什么:三个和稀疏无关的动机
SVT 的 CVar 描述把话说全了(SparseVolumeTextureTileDataTexture.cpp:14):
“Allocate the SVT tile data texture (streaming pool) as a reserved/virtual texture, backed by N small physical memory allocations to reduce fragmentation. This lifts the 2GB resource size limit and also allows for better GPU memory management when allocating the texture.”
| # | 动机 | 说明 |
|---|---|---|
| 1 | 抗碎片 | 256MB 一整块连续物理显存 → 16 个 16MB 小块。VSM 的注释:“This helps Windows video memory manager page allocations in and out of local memory more efficiently.” |
| 2 | 突破单资源尺寸上限 | SVT 平时受 MaxResourceSize = 2048MB 约束(SparseVolumeTextureUtility.h:15);开了 reserved 后上限改成 2048³ × 每体素字节数,只受 MaxVolumeTextureDim = 2048 这个维度限制(SparseVolumeTextureTileDataTexture.cpp:38-56) |
| 3 | 换页效率 | 物理内存分散成小块,OS 可以按块换入换出,而不是整块 256MB 一起搬 |
注意动机 2 的连带效应:它还带了个 r.SparseVolumeTexture.Streaming.ReservedResourcesMemoryLimit(默认 -1 = 不限)作为刹车,注释说 “Without this limit it is theoretically possible to allocate enormous amounts of memory.” —— 因为一旦解开 2GB 枷锁,就得自己防着分配爆炸。
❗ SVT 这条路径的门槛比其余用例都高一档:tile data texture 是 3D volume texture,FTileDataTexture::ShouldUseReservedResources() 查的是 GRHIGlobals.ReservedResources.SupportsVolumeTextures,也就是要求 Tiled Resources Tier 3,而不是全局那个 Tier 2 的 Supported。
5.6 真要改纹理大小:销毁重建
VSM 就是这么干的(VirtualShadowMapCacheManager.cpp:1100-1109):
if (!PhysicalPagePool
|| PhysicalPagePool->GetDesc().Extent != RequestedSize
|| PhysicalPagePool->GetDesc().ArraySize != RequestedArraySize
|| RequestedMaxPhysicalPages != MaxPhysicalPages
|| PhysicalPagePoolCreateFlags != RequestedCreateFlags)
{
if (PhysicalPagePool)
{
UE_LOGF(LogRenderer, Display,
"Recreating Shadow.Virtual.PhysicalPagePool due to size or flags change. This will also drop any cached pages.");
}
...
}整个池销毁重建 + 丢弃所有缓存页。 这也正好解释了为什么各家优化指南都说不要按关卡 / 按流送单元去改 r.Shadow.Virtual.MaxPhysicalPages —— 池重建不是免费的,reserved resource 一点也没让它变便宜(它优化的是分配方式,不是重建这件事本身)。
5.7 UE 里真正的「纹理稀疏」是软件虚拟化
| 系统 | 实现方式 |
|---|---|
| Virtual Texture (VT/RVT) | 自己维护 page table 纹理 + physical pool 纹理,shader 里做两次采样间接寻址 |
| Virtual Shadow Map | 同上:16K×16K 是概念上的虚拟分辨率,实际存储是固定大小的 physical page pool + page table |
| Sparse Volume Texture | 同上:page table volume + physical tile data volume |
三者都是把「稀疏」做在 shader 层,硬件 tiled resources 只被拿来当「抗碎片的分配器」。
⚠️ 原因没有写在任何注释里,但从代码分布看非常明显:
- 跨平台 —— 软件方案在 Metal / 移动端 / 主机上一致可用,硬件 tiled resources 的 Tier 支持参差不齐(UE 自己就把门槛卡在 Tier 2,3D 还要 Tier 3);
- 可控 —— page 大小、替换策略、反馈机制全在自己手里,不依赖驱动语义;
- 调试友好 —— page table 是一张普通纹理,RenderDoc 能直接看;而 tiled resources 的抓帧支持历史上一直有缺口。
六、UE5 里的三个消费者
三个消费者在 Windows 构建上全部处于开启状态,但动机各不相同:
| 系统 | 开关 | 默认状态 | 资源类型 |
|---|---|---|---|
| GPUScene | r.GPUScene.UseReservedResources | C++ 默认 true | Buffer(可伸缩) |
| Nanite Streaming | r.Nanite.Streaming.ReservedResources | BaseWindowsEngine.ini = 1(C++ 默认是 0) | Buffer(可伸缩) |
| VSM Page Pool | r.Shadow.Virtual.AllocatePagePoolAsReservedResource | C++ 默认 1 | Texture(全量提交) |
6.1 GPUScene:消灭 resize 时的 double buffer + memcpy
// Engine/Source/Runtime/Renderer/Private/GPUScene.cpp:151
static TAutoConsoleVariable<bool> CVarGPUSceneUseReservedResources(
TEXT("r.GPUScene.UseReservedResources"), true,
TEXT("Turning this on makes all GPU-Scene buffers default to using reserved resource allocations."), ECVF_ReadOnly);做法是一次性预留 2GB VA,之后只 commit 实际需要的部分(GPUScene.cpp:952):
constexpr uint64 ReservedAllocationSize = uint64(2048) << 20; // 2GB of address space
BufferDescReserved.Usage |= EBufferUsageFlags::ReservedResource;收益体现在这个分岔上(UnifiedBuffer.cpp:505):
if (EnumHasAllFlags(ExternalBuffer->Desc.Usage, EBufferUsageFlags::ReservedResource) && ...)
{
GraphBuilder.QueueCommitReservedBuffer(InternalBufferOld, BufferSizeNew);
return InternalBufferOld; // ← 同一个 buffer、同一个 GPU VA、descriptor 全部不失效
}
else
{
InternalBufferNew = GraphBuilder.CreateBuffer(...); // 新建
MemcpyResource(GraphBuilder, InternalBufferNew, InternalBufferOld, Params); // + 全量拷贝
return InternalBufferNew; // ← 峰值显存 = 旧 + 新,且所有 view 要重建
}Epic 在 5.8 的说明也直白确认了这个动机:“Enabled Reserved Resources for GPU Scene to reduce hitches on resize and peak GPU memory use by eliminating the need to double buffer during the copy.”
6.2 Nanite Streaming Pool:同样是「可增长池」
// Engine/Source/Runtime/Engine/Private/Rendering/NaniteStreamingManager.cpp:188
static int32 GNaniteStreamingReservedResources = 0; // ← C++ 默认值是 0
static FAutoConsoleVariableRef CVarNaniteStreamingReservedResources(
TEXT("r.Nanite.Streaming.ReservedResources"),
GNaniteStreamingReservedResources,
TEXT("Allow allocating Nanite GPU resources as reserved resources for better memory utilization and more efficient resizing (EXPERIMENTAL)"),
ECVF_ReadOnly | ECVF_RenderThreadSafe
);❗ 只看 C++ 默认值会得出错误结论。[SystemSettings] 会在启动时覆盖它:
; Engine/Config/Windows/BaseWindowsEngine.ini:23,在 [SystemSettings] 段内
r.Nanite.Streaming.ReservedResources=1这是上游 Epic 的平台默认设置,所以 Windows 上 Nanite 的 reserved resource 路径是开着的。
这条 ini 精确做了什么
| 维度 | 结论 |
|---|---|
| 生效范围 | Windows 平台的 Editor + Game 都吃;项目 DefaultEngine.ini 的 [SystemSettings] 可覆盖 |
| 可否运行时改 | ❌ CVar 标记为 ECVF_ReadOnly | ECVF_RenderThreadSafe。控制台改不动,只能改 ini 或启动加 -dpcvars=r.Nanite.Streaming.ReservedResources=0 |
| 只是「允许」,不是「强制」 | 实际生效还要 GRHIGlobals.ReservedResources.Supported:bReservedResource = (Supported && GNaniteStreamingReservedResources)(NaniteStreamingManager.cpp:550) |
| 不区分 RHI | D3D12 → Tiled Tier ≥ 2 才 true;Vulkan on Windows 也会命中(需 sparseBinding + sparseResidencyBuffer + sparseResidencyImage2D + residencyNonResidentStrict,且 graphics queue 有 VK_QUEUE_SPARSE_BINDING_BIT,且 SupportsParallelRendering());D3D11 下静默 fallback |
| 作用对象 | 只有 Nanite.StreamingManager.ClusterPageData 这一个 buffer。Hierarchy.DataBuffer 仍是普通 buffer(NaniteStreamingManager.cpp:573) |
⚠️ CVar 描述里写着 EXPERIMENTAL,但 Epic 已经在平台 ini 里默认打开了;标签看起来是历史遗留没清理,实际成熟度可参照 GPUScene(5.8 已在 release notes 里正式宣传)。
开启后 Nanite 内部行为的真实差异
对应 NaniteStreamingManager.cpp:1673-1833:
| 关闭(纯 C++ 默认值) | 开启(Windows 现状) | |
|---|---|---|
| buffer 创建 | CreateByteAddressDesc(4),靠 resize 长大 | 一次性预留 GetMaxPagePoolSizeInMB() 的 VA |
| root page 分配器 | FSpanAllocator(true) 强制 grow-only | 允许收缩(除非 r.Nanite.Streaming.Debug.ReservedResourceRootPageGrowOnly=1) |
| 初始 root page | NumInitialRootPages = 2048 → 2048×32KB = 64MB 起步 | 从 0 起步(ReservedResourceIgnoreInitialRootAllocation=1 默认开) |
| 增长粒度 | RoundUpToSignificantBits(N, 2),跳跃式 | 固定 16MB chunk(// Allocate pages in 16MB chunks to reduce the number of page table updates) |
| 池 resize 的实现 | 新建 buffer + AddCopyBufferPass 搬 root page,两份 allocation 同时存活 | 原地 AddPass_Memmove + commit/decommit,无临时峰值 |
| CSV 事件 | 发 GrowPoolAllocation | 不发(因为没有真正的 grow) |
Epic 自己在非 reserved 分支留的注释很说明问题:
// Non-reserved resource path: Make new allocation and copy root pages over.
// Temporary peak in memory usage when both allocations need to be live at the same time.
// ...
// It might not be worthwhile if reserved resources will be supported on all relevant platforms soon.而 reserved 分支要注意 memmove 的方向依赖 resize 方向(缩小时先搬后 resize,放大时先 resize 后搬)—— 因为 root page 就住在 streaming page 区域的上方,同一个 buffer 里,commit 是尾部伸缩的,顺序错了会踩到未提交的 tile。
一个调参上的连带变化:r.Nanite.Streaming.StreamingPoolSize(默认 512MB)现在改它不再触发 buffer 重建 + 全量拷贝,只是 memmove + 改页表。但仍然会重置流送状态 —— MaxStreamingPages 变化会走 bResetStreamingState 分支,UninstallAllResidentPages() 全部卸载并 memset 清零。所以「运行时调 pool size 会掉细节、要重新流送」这个现象依然存在,只是不再有显存尖峰和 hitch。
6.3 VSM Physical Page Pool(与 SVT):动机完全不同
// Engine/Source/Runtime/Renderer/Private/VirtualShadowMaps/VirtualShadowMapCacheManager.cpp:1094
// Using ReservedResource|ImmediateCommit flags hint to the RHI that the resource can be allocated using N small
// physical memory allocations, instead of a single large contighous allocation. This helps Windows video memory
// manager page allocations in and out of local memory more efficiently.
ETextureCreateFlags RequestedCreateFlags = (CVarVSMReservedResource.GetValueOnRenderThread() && GRHIGlobals.ReservedResources.Supported)
? (TexCreate_ReservedResource | TexCreate_ImmediateCommit) : TexCreate_None;CVar r.Shadow.Virtual.AllocatePagePoolAsReservedResource 默认 1。
注意它带了 ImmediateCommit —— 创建时就全量提交(D3D12Texture.cpp:538):
if (EnumHasAllFlags(Flags, TexCreate_ImmediateCommit))
{
// NOTE: Accessing the queue from this thread is OK, as D3D12 runtime acquires a lock around all command queue APIs.
Resource->CommitReservedResource(pDevice->GetQueue(ED3D12QueueType::Direct).D3DCommandQueue, UINT64_MAX /*commit entire resource*/);
}所以这里根本没用到「稀疏」这个能力,纯粹是借 reserved resource 把一块几百 MB 的大分配拆成 N 个 16MB heap,让 Windows 显存管理器换页更顺、避免大块连续显存申请失败。这是一个很值得学的「非典型用法」,动机见 §5.5。
同类还有 SparseVolumeTexture 的 tile data texture。有学术论文专门利用这点在 UE 里做超大体数据:“the backing memory can actually be allocated as a collection of many tiles… allowing a huge texture allocation even with fragmented VRAM”。
七、VA 账单与调优旋钮
7.1 一台 PC 实际预留了多少地址空间
Nanite 侧(NaniteStreamingManager.cpp:299):
static uint32 GetMaxPagePoolSizeInMB()
{
const uint32 DesiredSizeInMB = IsRHIDeviceAMD() ? 4095 : 2048;
const uint32 MaxSizeInMB = (uint32)(GRHIGlobals.MaxViewSizeBytesForNonTypedBuffer >> 20);
return FMath::Min(DesiredSizeInMB, MaxSizeInMB);
}D3D12RHI 没有覆写 MaxViewSizeBytesForNonTypedBuffer,走 RHI 默认 1ULL << 32 = 4096MB(RHIGlobals.h:324)。所以:
- NVIDIA / Intel → Nanite 预留 2048 MB VA
- AMD → 预留 4095 MB VA
再叠加 GPUScene(r.GPUScene.UseReservedResources C++ 默认就是 true,无需 ini),每个 reserved buffer 固定预留 2GB:
| Buffer | VA | 备注 |
|---|---|---|
GPUScene.PrimitiveData | 2 GB | |
GPUScene.InstanceSceneData | 2 GB | 仅 tiled 布局时(r.GPUScene.InstanceDataTileSizeLog2 默认 12;设为负数即关闭 tiled 布局并连带禁用 reserved) |
GPUScene.InstancePayloadData | 2 GB | |
GPUScene.LightmapData | 2 GB | |
Nanite.StreamingManager.ClusterPageData | 2 GB / 4 GB |
GPUScene.LightData 不在此列 —— 它走的是普通 buffer 路径。
合计约 10 GB(N 卡)~ 12 GB(A 卡)的 GPU 虚拟地址空间,物理显存占用是另一回事(只算 commit 的部分)。
对照告警阈值 rhi.ReservedResources.VirtualSizeWarningGB 默认 256GB(RHICoreStats.cpp:11)—— PC 上完全安全。⚠️ 这个默认值明显是给主机 / 受限 VA 平台留的余量,PC 上基本触发不了;真要触发说明有代码在循环创建 reserved 资源没释放,那时它是个很好的泄漏探针。
7.2 「16MB」这个数字为什么到处都是
GPUScene 的 tiled instance data 用了和 Nanite 一模一样的 16MB 对齐(GPUScene.cpp:988):
// Grow/shrink in 16MB chunks to reduce the number of page table updates or resizes
const uint32 AllocationStrideInBytes = 16u << 20;和 d3d12.ReservedResourceHeapSizeMB = 16 正好对齐 —— 一次 commit 恰好对应一个新 backing heap,不产生跨 heap 的碎片映射。这不是巧合。
背后的原则是 §2.5 那条:页表更新本身有成本,要按较粗粒度批量做。而且在 UE 的实现里还多两层代价 —— commit 会 CloseCommandList(),提交线程还要为它单独 Flush() 一次(§3.3)。Nanite + GPUScene 同帧各自 commit 一次,就会把该帧的提交批次切碎。这正是两边都用 16MB 粗粒度、且 Nanite 用 IgnoreInitialRootAllocation 从 0 起步按需长的原因:commit 次数比 commit 大小更值得优化。
7.3 backing heap 的对象数量
2048MB 池全提交时 = 128 个 16MB heap(Nanite 一家),加上 GPUScene 四个 buffer。这些 heap 都进 residency 管理的 handle 列表。
⚠️ 如果在 Insights 里看到 CommitReservedResource / residency MakeResident 的 CPU 时间偏高,d3d12.ReservedResourceHeapSizeMB 调到 32 或 64 是第一个该试的旋钮(ECVF_ReadOnly,需启动时设)。Vulkan 后端对应 r.Vulkan.SparseImageAllocSizeMB。
八、收益与代价清单
收益
- Resize 零拷贝:grow/shrink 只改页表,数据原地不动,GPU VA 不变 → 所有 SRV/UAV/descriptor 不失效(bindless 场景尤其值钱)。
- 削峰:省掉 resize 瞬间「旧 + 新」双份显存。
- 抗碎片:大资源由 N 个小 heap 拼成,不要求一块连续物理显存。
- 突破单资源尺寸上限:SVT 借此绕开 2GB
MaxResourceSize(§5.5)。 - VA 免费:预留 2GB 地址空间不花任何物理显存。
- 真稀疏(UE 目前基本没用):VT / SVT / 超大体纹理的天然载体。
代价与坑
- VA 是有预算的。UE 专门加了告警
rhi.ReservedResources.VirtualSizeWarningGB:“Some platforms have a limited virtual address space for reserved resources”。GPUScene 一个 buffer 就吃 2GB VA,多开几个要留意(§7.1)。 - ❗ 64KB 粒度取整:
AlignArbitrary(RequiredCommitSizeInBytes, 65536)—— 请求 1 字节实际占 1 个 tile(64KB),请求 65KB 实际占 2 个 tile(128KB)。小资源用它是净亏。 - packed mip 必须整体映射,且 array + 小 mip 在 Tier < 4 上不合法 → UE 用
TextureArrayMinimumMipDimension = 256保守规避。 - 页表更新有真实 CPU / 驱动开销,必须批量(Nanite 的 16MB 粒度就是这个原因);并且它会切断 command list,频繁 commit 会打碎提交批次。
- 部分队列不支持 → 触发跨队列 fence 串行化(§3.4)。
- 新提交区域内容未定义,必须先写再读。
- 纹理不能运行时伸缩,改尺寸只能销毁重建、丢掉全部缓存(§5.6)。
- 压缩(DCC)可能受影响,NVIDIA 建议按 2MB 对齐更新;UE 默认 16MB heap 天然满足,但单次 commit 粒度不一定。
- 工具链:RenderDoc 对 D3D12 tiled resources 的支持历史上就有缺口;PIX 对 reserved 资源是按 subresource 粒度做 active/inactive 追踪的,“inactive” 的 mip 内容不会被抓下来。抓帧调试要有心理预期。
九、如何验证它在工作
① 看已提交的物理量
stat nanitestreaming关注 Total Pool Size (MB) / Root Pool Size (MB) / Streaming Pool Size (MB) —— 这些是已提交的物理量。
② 看预留 vs 实际的差值
stat rhiEngine/Source/Runtime/RHICore/Private/RHICoreStats.cpp 提供的口径:
STAT_ReservedUncommittedTextureMemory/STAT_ReservedUncommittedBufferMemory= 预留但未提交(buffer 侧应该很大,≈10GB 量级)STAT_ReservedCommittedBufferMemory/STAT_ReservedCommittedTextureMemory= 已提交STAT_D3D12ReservedResourcePhysical= 实际 backing heap 字节数(含 slack,所以可能略大于「已提交」)GRHIGlobals.ReservedResources.VirtualSize= 所有 reserved 资源的 VA 总量
Uncommitted 与 Committed 之和恒等于预留总量(UpdateReservedResourceStatsOnCommit 是一 INC 一 DEC)。Uncommitted 这一项,就是这个特性替你省下的显存。
③ 判定是否真的走了 reserved 路径
最直接的两个办法:
- 在 PIX / Insights 里找
CommitReservedResource这个TRACE_CPUPROFILER_EVENT_SCOPE; - 看是否还有
CSV_EVENT(NaniteStreaming, "GrowPoolAllocation")—— 走 reserved 路径时这个事件永远不会出现。
十、总结
10.1 从 UE4 心智模型迁移
UE4 你只有 「资源 = 一块内存」。DX12 把它拆成 「资源 = 一段地址 + 一张页表 + 一堆物理页」:
- Committed = 地址和页由运行时打包给你,永久绑定;
- Placed = 页由你管(heap),但绑定关系创建时定死、且必须连续;
- Reserved = 连绑定关系都交给你,运行时可改、可空、可来自多个 heap、可多对一。
而 UE5 目前主要吃的不是它的「稀疏」能力,而是**「地址不变、物理内存可原地伸缩」**这一条 —— 这正好治好了 GPU-driven 渲染里最痛的那个病:大 buffer 扩容要 double buffer + 全量 memcpy + 重建所有 view。
10.2 Buffer vs Texture 一页对照
| Reserved Buffer | Reserved Texture | |
|---|---|---|
| UE 是否支持部分提交 | ✅ 支持,QueueCommitReservedBuffer | ❌ 不支持,validation 直接报错 |
| 提交时机 | 运行时任意次,尾部伸缩 | 仅创建时一次(ImmediateCommit 全提交) |
| 典型使用者 | GPUScene 四大 buffer、Nanite ClusterPageData | VSM PagePool、SVT TileData |
| 核心收益 | resize 零拷贝、descriptor 不失效、无双份峰值 | 抗碎片、突破 2GB 单资源上限、换页效率 |
| 为什么能 / 不能伸缩 | 一维无结构,尾部追加语义天然成立 | 多维有结构;且 mip0 在 tile 空间最前面,尾部水位线砍不到 |
| 想「变大」怎么办 | 直接 commit 更多 tile | 销毁重建(VSM 就是这么做的,会丢缓存) |
| 硬件「真稀疏」用了吗 | 没有(线性水位线) | 没有(全提交) |
一句话:UE5 把 DX12 reserved resource 当成了两样东西 —— 对 buffer,它是「地址不变的可伸缩显存水位阀」;对 texture,它是「把大块分配拆成小块的抗碎片分配器」。两者都没有用到它最出名的那个能力:稀疏驻留。
十一、参考链接
官方规格
- 🔗 Memory Management Strategies — 三种资源类型的权威定义
- 🔗 Residency — reserved 的驻留语义
- 🔗 D3D12_TILED_RESOURCES_TIER — Tier 1~4 逐条能力,最该精读的一页
- 🔗 ID3D12CommandQueue::UpdateTileMappings — 四种 range flag 的完整语义与代码示例
- 🔗 D3D12_TEXTURE_LAYOUT — 64KB_UNDEFINED / STANDARD_SWIZZLE
- 🔗 Volume Tiled Resources — 3D 标准 tile 形状表
- 🔗 ID3D12GraphicsCommandList::CopyTiles — 明确说明 packed mip 不能用 CopyTiles
- 🔗 D3D12 Tiled Resource Tier 4 — DirectX-Specs — array + full mip chain 限制的来龙去脉
- 🔗 Porting from D3D11 to D3D12 — tile pool → heap 的映射关系
- 🔗 Tiled resources (D3D11) — 概念起源
厂商 / 工具
- 🔗 Advanced API Performance: Memory and Resources — NVIDIA ⭐ 最实用的性能 do/don’t 清单
- 🔗 GPU Captures: placed and reserved resources — PIX on Windows ⭐ 抓帧行为与限制
- 🔗 renderdoc#2203 — Support for Tiled Resources in D3D12
教程 / 示例
- 🔗 Resource Handling in D3D12 — Riccardo Loggini ⭐ 三类资源 + 对齐 + aliasing barrier 讲得最系统
- 🔗 DirectX-Graphics-Samples / D3D12ReservedResources — 官方最小可运行示例(按需 map/unmap mip)
- 🔗 Differences in memory management between Vulkan and D3D12 — Adam Sawicki — 对照 Vulkan sparse binding
- 🔗 [D3D12] Minimal Tiled Resources implementation — GameDev.net — 踩坑贴
UE5 相关
- 🔗 Unreal Engine 5.8 Performance Highlights — Tom Looman ⭐ 明确 GPU Scene 启用 reserved resources 的动机
- 🔗 r.Nanite.Streaming.ReservedResources — Unreal Directive
- 🔗 Rendering Large Volume Datasets in UE5: A Survey (arXiv 2504.07485) ⭐ 用 reserved resources 突破 1 gigavoxel 限制的实证