来源:DX12 引擎源码
引擎版本:UE 5.8
面向读者:UE 图形程序员


DX12 里 committed / placed / reserved 三种资源,前两种在 UE 里天天见,reserved 却直到 5.8 才在 PC 上默认跑起来 —— GPUScene 的四个大 buffer、Nanite 的 ClusterPageData、VSM 的 physical page pool,全都是它的消费者。

有意思的是:UE 用它,用的几乎都不是它最出名的那个能力(稀疏驻留)。

目录

  1. 三种资源,同一张坐标系
  2. 硬件与 API 层原理
  3. UE5 的接入面
  4. 核心实现精读:CommitReservedResource
  5. Buffer 与 Texture 的分野
  6. UE5 里的三个消费者
  7. VA 账单与调优旋钮
  8. 收益与代价清单
  9. 如何验证它在工作
  10. 总结
  11. 参考链接

一、三种资源,同一张坐标系

DX12 资源 = GPU 虚拟地址空间(VA) + 物理页(heap) + 两者之间的映射(页表)。DX11 把这三者绑死成一个东西,DX12 才把它们拆开:

VA 谁分配物理内存谁分配映射关系能否改映射
Committed运行时隐式运行时隐式(隐式 heap,刚好装下资源)创建时 1:1 固定❌
Placed运行时(落在 heap 的 VA 里)先 CreateHeap,再给 offset创建时固定,必须连续❌(要换只能销毁重建 + 重建 descriptor)
Reserved只分配 VA,全部页初始为 NULL完全由你控制,可来自多个 heap运行时任意改,可稀疏✅ UpdateTileMappings

❗ 一个命名陷阱先说在前面:UE 里 CommitReservedResource() 的 “commit”,指的是给 reserved 资源的一段虚拟范围建立物理 backing,不是把它变成上表里的 Committed Resource。资源从 CreateReservedResource() 创建的那一刻起就永远是 reserved,不存在「转化」这回事。

微软对 reserved 的定义是:它拥有自己独立的 GPU 虚拟地址空间,允许先大量预留 VA,之后再把 VA 页映射到 heap 的某些区域,并且随时可以重新配置;当一段 VA 映射到 NULL 或未驻留的 heap 时,该部分资源被视为 non-resident。

PIX 团队给过一个很好的一句话类比:placed resource 在概念上就是一个只有一条、且永远不能改的映射的 reserved resource。

跨 API 对照:Reserved 就是 D3D11 的 Tiled Resources,也就是 Vulkan 的 Sparse Binding / Sparse Residency、OpenGL 的 Partially Resident Textures。


二、硬件与 API 层原理

2.1 Tile:映射的最小单位是 64KB

Reserved Resource 的 VA 空间(连续、巨大、免费)
┌──────┬──────┬──────┬──────┬──────┬──────┬──────┬──────┐
│ T0   │ T1   │ T2   │ T3   │ T4   │ T5   │ T6   │ T7   │  每格 = 64KB tile
└──┬───┴──┬───┴──┬───┴──┬───┴──────┴──────┴──────┴──────┘
   │      │      │      │        ↑ 这些 tile 映射到 NULL(不占物理显存)
   ▼      ▼      ▼      ▼
┌────────────────┐  ┌────────────────┐
│  ID3D12Heap A  │  │  ID3D12Heap B  │   ← 物理页可来自多个 heap(D3D12 相对 D3D11 的关键放宽)
└────────────────┘  └────────────────┘
  • Buffer:64KB 线性切分,GetResourceTiling 返回一个平凡的 1D subresource tiling。
  • Texture:tile 形状是「各维度近似等长的 64KB 区域」,这是 D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE 的定义要求。

3D(Volume)的标准 tile 形状官方给了表:8bpp→64×32×32、16bpp→32×32×32、32bpp→32×32×16、64bpp→32×16×16、128bpp→16×16×16、BC1/4→128×64×16、BC2/3/5/6/7→64×64×16。

2D 情形同理由 64KB 约束推出:8bpp 256×256、16bpp 256×128、32bpp 128×128、64bpp 128×64、128bpp 64×64 —— 每项乘起来正好 65536 字节。

⚠️ 这组 2D 数值没有对应的官方表页面,是从「64KB + 近似等边」两条约束反推的业界通行值。工程上不要硬编码,运行时以 GetResourceTiling 返回的 D3D12_TILE_SHAPE 为准(UE 就是这么做的)。

分层视角:一次访问要穿过五层

上面那张图只画了「VA tile ↔ heap tile」这一层。完整链路有五层,混淆它们是理解 reserved resource 时最常见的卡点:

Reserved resource 的分层内存模型:tile mapping 层与 residency 层是两件独立的事

关键在于 tile mapping 层和 residency 层相互独立:一个 resource tile 已经指向某个 heap tile(映射建好了),不代表那个 heap 此刻驻留在 GPU 可访问的内存里。UE 分别用 UpdateTileMappings 和 Residency Manager 管这两层,两者之间的时序约束见 §4.4 阶段 6。

2.2 Packed Mip(mip tail)—— 最容易踩坑的地方

当某级 mip 的尺寸小于一个 tile 时,硬件会把剩下的所有小 mip 打包进 1~N 个 tile,称为 packed mips / mip tail。GetResourceTiling 通过 D3D12_PACKED_MIP_INFO 告诉你 NumStandardMips、NumPackedMips、NumTilesForPackedMips。

约束是硬的:

  • packed mip 必须整体一次性映射完,不能只映射其中一部分(UE 代码里对此有专门的 assert,见 §4.5);
  • Tier 2/3 下,「数组切片 > 1」且「存在小于 tile 的 mip」的组合不允许;Tier 4 才放开这条限制。

2.3 Tier 分级:决定你能不能用、用得多放心

Tier能力说明
NOT_SUPPORTEDCreateReservedResource 完全不可用,连 buffer 都不行
Tier 1可创建 2D reserved 纹理与 buffer⚠️ 读写 NULL tile 是 undefined;无 LOD clamp 指令;官方建议把同一个 page 重复映射到所有「本该 NULL」的位置来规避
Tier 2mip 组织有明确保证;读 NULL tile 返回 0,写 NULL tile 被丢弃;shader 提供 LOD clamp 与 residency 反馈(Sample(S,float,int,float,uint) + CheckAccessFullyMapped)Feature Level 12_0 的适配器全部保证 ≥ Tier 2
Tier 3增加 3D(Volume)Tiled Resources
Tier 4数组纹理可带完整 mip 链(含 packed mip)

2.4 API 面:比 placed 多出来的那一套

API作用执行位置
ID3D12Device::CreateReservedResource(1)只建 VA,无 backing store立即(CPU)
ID3D12Device::GetResourceTiling查询 tile 数、tile 形状、packed mip 信息立即(CPU)
ID3D12CommandQueue::UpdateTileMappings改页表:把 tile 区间映射到 heap 区间 / NULL / 单 tile 复用 / 跳过Queue 级操作,不是 command list 操作
ID3D12CommandQueue::CopyTileMappings把一个 reserved 资源的映射整体拷到另一个Queue 级
ID3D12GraphicsCommandList::CopyTiles在 tile 与线性 buffer 之间搬数据(不是映射)Command list

UpdateTileMappings 的 range flag 四选一:NONE(顺序映射一段 heap tile)、REUSE_SINGLE_TILE(多个 VA tile 共享同一物理 tile,这就是 Tier 1 规避 NULL 的手段)、NULL(解映射)、SKIP(保持原样)。

❗「Queue 级」这三个字是理解 UE 实现的钥匙:它按队列提交顺序执行,不进 command list,所以引擎必须把它安排在正确的提交时点上,并且跨队列时要显式 fence 同步(见 §3.3、§3.4)。

2.5 NVIDIA 的实战建议

  • 把所有 UpdateTileMappings 移到异步 copy queue,用来隐藏 OS 调度与提交开销;
  • 只要是重新映射同一批 tile,就不需要显式 unmap;
  • 避免用 pHeap = NULL / D3D12_TILE_RANGE_FLAG_NULL 做显式 unmap —— 那会强制驱动遍历所有映射并移除不再映射的 tile;宁可切到另一个 tile;
  • 按 2MB 粒度对齐 tile 更新,以便获得最优的硬件压缩(DCC 等)。

其中第一条和 UE 的选择是相反方向的取舍,见 §3.4。


三、UE5 的接入面

3.1 能力检测:门槛定在 Tier 2

// Engine/Source/Runtime/D3D12RHI/Private/Windows/WindowsD3D12Device.cpp:1904
// Tier 2 is guaranteed for all adapters with feature level 12_0.
GRHIGlobals.ReservedResources.Supported              = Options.TiledResourcesTier >= D3D12_TILED_RESOURCES_TIER_2;
// Tier 3 is required to create volume textures. Some hardware may support it.
GRHIGlobals.ReservedResources.SupportsVolumeTextures = Options.TiledResourcesTier >= D3D12_TILED_RESOURCES_TIER_3;

UE 把门槛直接定在 Tier 2,而不是 API 最低要求的 Tier 1 —— 因为 Tier 1 的「读 NULL tile 未定义」对引擎来说没法安全使用。

RHI 全局能力块在 RHIGlobals.h:720-751:

struct FReservedResources
{
    bool  Supported = false;
    bool  SupportsVolumeTextures = false;
    int32 TextureArrayMinimumMipDimension = 256;    // 规避 §2.2 的 array + 小 mip 限制(保守、与格式无关)
    static constexpr int32 TileSizeInBytes = 65536; // 跨平台统一为 64KB
    volatile int64 VirtualSize = 0;                 // 所有 reserved 资源占用的 VA 总量
} ReservedResources;

TextureArrayMinimumMipDimension = 256 这个保守值,就是为了不去碰「Tier 4 才放开的那条限制」。

3.2 创建侧:两个新 flag 与几条硬约束

RHIDefinitions.h:966 / :1171:

  • EBufferUsageFlags::ReservedResource(BUF_ReservedResource)
  • ETextureCreateFlags::ReservedResource(TexCreate_ReservedResource)
  • ETextureCreateFlags::ImmediateCommit(TexCreate_ImmediateCommit)—— 创建即全量提交

注释里明确写了 EXPERIMENTAL,且不可与 Dynamic 等「不能落在 local GPU memory」的 flag 共用。

创建路径 D3D12Resources.cpp:887 里的硬性 check 正好对应 §2.1。注意 layout 那条只对 1D/2D/3D 纹理生效,buffer 不受此约束:

// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp
if (LocalDesc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE1D
    || LocalDesc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE2D
    || LocalDesc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE3D)
{
    checkf(LocalDesc.Layout == D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE, ...);
}
checkf(LocalDesc.Alignment == 0 || LocalDesc.Alignment == 65536, ...);
// ...
// NOTE: reserved resource residency is not tracked/managed by the engine, so we don't need to call StartTrackingForResidency().

最后那行 NOTE 很重要:reserved 资源本身不做 residency 追踪,驻留单位是 backing heap(见 §4.4 阶段 6)。

Buffer 侧的约束在 D3D12Buffer.cpp:486:

checkf(!bHasInitialData,   TEXT("Reserved resources may not have initial data"));
checkf(!bIsDynamic,        TEXT("Reserved resources may not be dynamic"));
checkf(!ResourceAllocator, TEXT("Reserved resources may not use a custom resource allocator"));

3.3 提交侧:UE 把 commit 挂在了 Transition 上

这是一个很「UE」的设计选择 —— 不新增一条 RHI 命令,而是复用资源转换:

// Engine/Source/Runtime/RHI/Public/RHITransition.h:103-116
/**
* Represents a change in physical memory allocation for a resource that was created with TexCreate/BUF_ReservedResource flag.
* Physical memory is allocated in tiles/pages and mapped to the tail of the currently committed region of the resource.
* This API may be used to grow or shrink reserved resources without moving the bulk of the data or re-creating SRVs/UAVs.
* The contents of the newly committed region of the resource is undefined and must be overwritten by the application before use.
* Reserved resources must be created with maximum expected size, which will not cost any memory until committed.
* Commit size must be smaller or equal to the maximum resource size specified at creation.
* Check GRHIGlobals.ReservedResources.Supported before using this API or TexCreate/BUF_ReservedResource flag.
*/
struct FRHICommitResourceInfo
{
    uint64 SizeInBytes = 0;
    explicit FRHICommitResourceInfo(uint64 InSizeInBytes) : SizeInBytes(InSizeInBytes) {}
};

调用链(自上而下):

FRDGBuilder::QueueCommitReservedBuffer(Buffer, NewSize)             RenderGraphBuilder.inl:427
   └─> FRDGPooledBuffer::SetCommittedSize()                          RenderGraphResources.h:1273
        └─> FRHITransitionInfo{ Buffer, Before, After, FRHICommitResourceInfo }
             └─> FD3D12ContextCommon::SetReservedBufferCommitSize()   D3D12CommandContext.cpp:332
                  ├─ if (IsPendingCommands()) CloseCommandList()   ← 先关掉当前 command list
                  └─ GetPayload(EPhase::UpdateReservedResources)->ReservedResourcesToCommit.Add(...)
                       └─> FD3D12DynamicRHI::UpdateReservedResources(Payload)   D3D12Submission.cpp:649
                            └─> FD3D12Resource::CommitReservedResource(Queue, Size)  D3D12Resources.cpp:223

这里的 CloseCommandList() 和独立的 EPhase::UpdateReservedResources payload 阶段,正是 §2.4 那句「queue 级操作」在架构上的落地:页表更新必须切在两批 command list 之间,不能混在里面。

而且切断发生在两个层次上 —— 除了上面这条 command list 的关闭,提交线程在遍历 payload 时也会为它单独 flush 一次:

// Engine/Source/Runtime/D3D12RHI/Private/D3D12Submission.cpp:887
if (Payload->HasUpdateReservedResourcesWork())
{
    Flush();
    UpdateReservedResources(Payload);
}

这条性质的连带成本见 §8。

3.4 队列能力兜底

// Engine/Source/Runtime/D3D12RHI/Private/D3D12Submission.cpp:653
// On some devices, some queues cannot perform tile remapping operations.
// We can work around this limitation by running the remapping in lockstep on another queue:
// - tile mapping queue waits for commands on this queue to finish
// - tile mapping queue performs the commit/decommit operations
// - this queue waits for tile mapping queue to finish
ID3D12CommandQueue* TileMappingQueue = (Queue.bSupportsTileMapping ? Queue.D3DCommandQueue : Queue.Device->TileMappingQueue).GetReference();
const bool bCrossQueueSyncRequired = TileMappingQueue != Queue.D3DCommandQueue.GetReference();

不支持 tile mapping 的队列(如某些硬件上的 copy / compute queue)会回退到 Direct queue(D3D12Device.cpp:135),代价是一次双向 fence 的串行化:

跨队列 tile mapping 的同步时序:两次等待缺一不可

前一次等待防止「旧访问还没做完就改页表」,后一次防止「新映射还没生效就访问新区域」。

⚠️ 这与 NVIDIA 建议的「把 UpdateTileMappings 放到 async copy queue」是相反方向的取舍 —— UE 选择跟随提交队列、必要时兜底,而不是无脑丢给 copy queue,因为 RDG 的 commit 时点与 pass 依赖强绑定,跨队列会引入额外的同步点。


四、核心实现精读:CommitReservedResource

实现在 Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp:223-611。

4.1 心智模型:一条只在尾部伸缩的水位线

这个函数只干一件事:

把「我要多少物理显存」这个字节数,翻译成一串 UpdateTileMappings 调用,并顺便管理背后的 heap。

对上 C++ 容器的直觉:

CreateReservedResource(MaxSize)  ≈  reserve():确定 capacity,此后不可改
CommitReservedResource(NewSize)  ≈  resize() :改变当前「有物理 backing」的 size

类比到此为止 —— 资源本身永远是 reserved resource,不会被「转化」成 committed resource;而且 resize() 会保留旧数据,Commit 新增出来的区域内容是未定义的。

关键设计:UE 的 reserved resource 不是「任意稀疏」,而是「线性尾部伸缩的池」。

Reserved Resource 的 tile 空间(VA,创建时就固定,比如 32768 个 tile = 2GB)
┌─────────────────────────────────────────────────────────────────────┐
│■■■■■■■■■■■■■■■■■■■■■│░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░│
└─────────────────────┬───────────────────────────────────────────────┘
  已提交(有物理内存)  │  未提交(映射到 NULL,不占显存)
  NumCommittedTiles ───┘  ← 这个函数就是在左右移动这条线

没有「中间挖洞」,没有随机稀疏,只有一个标量 NumCommittedTiles。想通这一点,函数就塌掉一半难度 —— 这也解释了为什么它主要服务于「可增长的大 buffer」,而不是「virtual texture 式的随机稀疏」。

物理内存那边长这样:

NumCommittedTiles = 640
 tile:  [0 ────────── 255][256 ───────── 511][512 ─── 639]
         └── Heap A ──┘    └── Heap B ──┘    └─ Heap C ─┘
           16MB/256格        16MB/256格        8MB/128格
                                              ↑ BackingHeaps 数组,顺序 = tile 顺序

后文会把这套映射关系写成「概念页表」的形式:

PTE[R[i]] = { Heap 对象, Heap 内 tile 序号 }     或     PTE[R[i]] = NULL

这只是理解 API 语义的模型,不是 D3D12 暴露给应用的真实硬件页表结构。真实实现里 GPU 页表可能是多级的、heap 对应的 WDDM allocation 可能被分页、实际页面可能在本地 VRAM 也可能被 OS 迁移。D3D12 向应用暴露的稳定抽象只有一条:reserved resource 的 tile ↔ 某个 ID3D12Heap 里的 tile offset。

4.2 四个状态变量与六条不变量

// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.h:223
struct FD3D12ReservedResourceData
{
    TArray<TRefCountPtr<FD3D12Heap>> BackingHeaps;
 
    // Flattened array of residency handles owned by backing heaps, used to support batched GetResidencyHandles()
    TArray<FD3D12ResidencyHandle*> ResidencyHandles;
    TArray<int32> NumResidencyHandlesPerHeap;
 
    // Tiles currently assigned to the resource
    uint32 NumCommittedTiles = 0;
 
    // Available tiles at the end of the last backing heap
    uint32 NumSlackTiles = 0;
};
含义一句话
BackingHeaps物理内存的分段列表,顺序即 tile 顺序第 i 个 heap 服务 tile 区间的第 i 段
ResidencyHandles拍平的驻留句柄,NumResidencyHandlesPerHeap 记录每个 heap 贡献了几个,方便 pop 时同步弹出记账用
NumCommittedTiles水位线。tile [0, N) 有内存全函数的核心状态
NumSlackTiles最后一个 heap 尾部有多少 tile 已解映射但 heap 还留着见下

❗ NumSlackTiles 是唯一反直觉的地方。新建 heap 时,尺寸是按需精确分配的:

const uint32 ThisHeapSize = RegionSize.NumTiles * TileSizeInBytes;   // 要几个 tile 就建多大

所以新建的 heap 永远没有 slack。NumSlackTiles > 0 只可能由「收缩」产生:收缩时如果某个 heap 只被解映射了一部分,这个 heap 不能删(前半段还在用),于是它尾部就空出来一段 —— 记成 slack,下次增长优先复用,避免「删了又建」的抖动。

收缩前:  [Heap B: ████████████████]  256 tiles 全用
收缩后:  [Heap B: █████░░░░░░░░░░░]  用了 44,slack = 212(heap 还在,内存还占着)
再增长:  [Heap B: █████████░░░░░░░]  从 offset 44 复用 100 个,slack = 112

必须维持的六条不变量

整个函数的读法可以简化成「检查这六条永远成立」:

  1. 已映射区域始终是 [0, NumCommittedTiles);
  2. 资源中间不留洞;
  3. BackingHeaps 按资源虚拟 tile 顺序排列(第 i 个 heap 服务第 i 段);
  4. 只有最后一个 heap 可以含 slack;
  5. 扩容和缩容都只从尾部进行;
  6. 保留下来的前缀仍指向原来的 heap tile,因此永远不需要复制数据。

第 3、4 条是 §4.4 那两段 while 循环能成立的全部前提 —— 正因为 slack 只可能出现在最后一个 heap,HeapFirstTile = NumCommittedTiles - NumUsedTilesInHeap 这种「从水位线往回倒推 heap 起点」的算法才是对的。

由此得到两个记账公式:

CommittedBytes    = NumCommittedTiles × 64KB
PhysicalHeapBytes = CommittedBytes + NumSlackTiles × 64KB

第二个公式解释了一个实际现象:stat rhi 里的 STAT_D3D12ReservedResourcePhysical(backing heap 实际字节数)可能大于逻辑上已提交的字节数,差额就是 slack。这个统计量在 CreateHeap 时 INC、在 DeferDelete() 时 DEC,跟的是 heap 的生死,而不是水位线。

4.3 函数骨架:六个阶段

void FD3D12Resource::CommitReservedResource(ID3D12CommandQueue* D3DCommandQueue, uint64 RequiredCommitSizeInBytes)
{
    // 阶段 1:前置检查
    // 阶段 2:问 D3D「这个资源是怎么切 tile 的」
    // 阶段 3:算账(目标 tile 数、heap 尺寸上限)
    // 阶段 4:定义「线性 tile 序号 → D3D 坐标」的换算 lambda
    // 阶段 5:走 收缩 或 增长 分支,产出 MappingParams 数组(此时还没调 D3D)
    // 阶段 6:MakeResident → 批量 UpdateTileMappings → SignalFence → 更新统计
}

画成三段结构(阶段 4 只是定义一个 lambda,不占流程上的位置,故未画出):

阶段 5 只往数组里记条目,一次 D3D 都不调;真正的执行集中在阶段 6

5A 和 5B 会按 heap 逐轮循环(一次收缩或增长可能跨越多个 heap),两条分支最终汇合到同一处,MakeResident → UpdateTileMappings → SignalFence 只在阶段 6 发生一次。

之所以只「记账」不「执行」:MakeResident 必须发生在 UpdateTileMappings 之前(不能把 tile 映射到一个还没驻留的 heap),而新 heap 是在阶段 5 里才创建的。所以必须先全部算完、收集齐 residency handle,再统一执行。

4.4 逐阶段精读

阶段 1:前置检查(L225-242)

TRACE_CPUPROFILER_EVENT_SCOPE(CommitReservedResource);
 
static constexpr uint64 TileSizeInBytes = GRHIGlobals.ReservedResources.TileSizeInBytes;
static_assert(TileSizeInBytes == 65536, "Reserved resource tiles are expected to always be 64KB");
 
check(Desc.bReservedResource);
check(ReservedResourceData.IsValid());
LLM_REALLOC_SCOPE(ReservedResourceData.Get());   // 内存统计:把后面 CreateHeap 的分配挂到这个资源名下
UE_MEMSCOPE_PTR(ReservedResourceData.Get());
 
checkf(GRHIGlobals.ReservedResources.Supported, ...);
if (Desc.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE3D)
    checkf(GRHIGlobals.ReservedResources.SupportsVolumeTextures, ...);   // 3D 需要 Tier 3

LLM_REALLOC_SCOPE(ReservedResourceData.Get()) 用这个结构体的地址当作 LLM 的分配 key —— 因为 reserved 资源的「地址」会随着 heap 增删而变,没法用 GPU VA 当 key。

顺带一提,函数开头那个 TRACE_CPUPROFILER_EVENT_SCOPE(CommitReservedResource) 就是后面 §9 里在 Insights 判定「是否真的走了 reserved 路径」的抓手。

阶段 2:问 D3D 资源怎么切 tile(L244-272)

uint32 D3DResourceNumTiles = 0;              // out: 整个资源一共多少 tile
D3D12_PACKED_MIP_INFO PackedMipDesc = {};    // out: 标准 mip / 打包 mip 的分界
D3D12_TILE_SHAPE TileShape = {};             // out: 一个 tile 是多少 texel(这里查了但没用)
uint32 NumSubresourceTilings = NumMipLevels; // in/out
TArray<D3D12_SUBRESOURCE_TILING, TInlineAllocator<16>> MipTilingInfo;
MipTilingInfo.SetNum(NumSubresourceTilings);
 
D3DDevice->GetResourceTiling(GetResource(), &D3DResourceNumTiles, &PackedMipDesc,
                             &TileShape, &NumSubresourceTilings, /*FirstSubresource*/0,
                             MipTilingInfo.GetData());

两个返回结构:

struct D3D12_PACKED_MIP_INFO {
    UINT8 NumStandardMips;        // 前 N 级 mip 是「标准 tile 排布」
    UINT8 NumPackedMips;          // 后面这些小 mip 被打包进 mip tail
    UINT  NumTilesForPackedMips;  // mip tail 占几个 tile(每个 array slice)
    UINT  StartTileIndexInOverallResource;
};
struct D3D12_SUBRESOURCE_TILING {   // 每级标准 mip 一个
    UINT   WidthInTiles, StartTileIndexInOverallResource;
    UINT16 HeightInTiles, DepthInTiles;
};

这里体现了 §2.1 那条「不要硬编码 tile 形状」的原则:一切以运行时查询为准。另外注意它只查了 array slice 0 的 mip 链(FirstSubresource = 0,数量 = NumMipLevels)—— 因为所有 slice 的 tiling 完全一样,后面用乘法补回 slice 维度。

然后是那个「作弊」:

if (bBuffer)
{
    // Buffers obviously don't have mips, but we can pretend they do to make the code below agnostic to resource type
    PackedMipDesc.NumStandardMips = 1;
}
check(MipTilingInfo.Num() == PackedMipDesc.NumStandardMips + PackedMipDesc.NumPackedMips);

buffer 没有 mip,D3D 返回的 PackedMipDesc 全 0。UE 强行填 NumStandardMips = 1,把 buffer 伪装成「1 级标准 mip 的纹理」,这样下面所有代码只有一套逻辑,不需要 if/else 分叉。

阶段 3 里还有一条与之配套的断言,把这个伪装的前提钉死:

checkf(D3DResourceNumTiles == MipTilingInfo[0].WidthInTiles,
    TEXT("Reserved buffers are expected to have trivial tiling configuration: single 1D subresource that contains all tiles."));

阶段 3:算账(L274-324)

const uint32 NumPackedTilesPerArraySlice = PackedMipDesc.NumTilesForPackedMips;
const uint32 NumTotalPackedMipTiles      = NumPackedTilesPerArraySlice * NumArraySlices;
const uint32 NumTotalStandardMipTiles    = D3DResourceNumTiles - NumTotalPackedMipTiles;
 
const uint64 TotalSize = D3DResourceNumTiles * TileSizeInBytes;
 
RequiredCommitSizeInBytes = FMath::Min<uint64>(RequiredCommitSizeInBytes, TotalSize);   // 传 UINT64_MAX = 全提交
RequiredCommitSizeInBytes = AlignArbitrary(RequiredCommitSizeInBytes, TileSizeInBytes); // 上取整到 64KB
 
const uint64 MaxHeapSize     = uint64(CVarD3D12ReservedResourceHeapSizeMB.GetValueOnAnyThread()) * 1024 * 1024;  // 默认 16MB
const uint64 NumHeaps        = FMath::DivideAndRoundUp(TotalSize, MaxHeapSize);    // 仅用于 BackingHeaps.Reserve()
const uint32 MaxTilesPerHeap = uint32(MaxHeapSize / TileSizeInBytes);              // 默认 256
 
const uint32 NumRequiredCommitTiles = RequiredCommitSizeInBytes / TileSizeInBytes; // ★ 目标水位线

UINT64_MAX 是「全部提交」的约定值(TexCreate_ImmediateCommit 就走这条)。

物理内存不是一个大 heap,而是一串默认 16MB 的小 heap:

// Engine/Source/Runtime/D3D12RHI/Private/D3D12Resources.cpp:16
static TAutoConsoleVariable<int32> CVarD3D12ReservedResourceHeapSizeMB(
    TEXT("d3d12.ReservedResourceHeapSizeMB"), 16,
    TEXT("Size of the backing heaps for reserved resources in megabytes (default 16MB)."), ECVF_ReadOnly);

Vulkan 后端有一个一模一样的旋钮 r.Vulkan.SparseImageAllocSizeMB(VulkanTexture.cpp:43),默认值同样是 16MB,同样 ECVF_ReadOnly。

heap 的 residency 优先级沿用 committed resource 的启发式,高优先级的会额外调一次 SetResidencyPriority(D3D12_RESIDENCY_PRIORITY_HIGH):

const bool bRenderOrDepthTarget  = Flags & (ALLOW_RENDER_TARGET | ALLOW_DEPTH_STENCIL);
const bool bHighPriorityResource = bRenderOrDepthTarget || (Flags & ALLOW_UNORDERED_ACCESS);

heap flag 也按用途收紧(这会影响某些硬件的压缩 / 对齐行为):

const D3D12_HEAP_FLAGS TextureHeapFlags = bRenderOrDepthTarget
    ? D3D12_HEAP_FLAG_ALLOW_ONLY_RT_DS_TEXTURES
    : D3D12_HEAP_FLAG_ALLOW_ONLY_NON_RT_DS_TEXTURES;
const D3D12_HEAP_FLAGS HeapFlags = bBuffer ? D3D12_HEAP_FLAG_ALLOW_ONLY_BUFFERS : TextureHeapFlags;

阶段 5A:收缩分支(L395-453)

if (ReservedResourceData->NumCommittedTiles > NumRequiredCommitTiles)   // 要变小
{
    while (NumCommittedTiles > NumRequiredCommitTiles)
    {
        // ── 1. 定位「最后一个 heap」覆盖的 tile 区间 ──
        TRefCountPtr<FD3D12Heap>& LastHeap = BackingHeaps.Last();
        const uint32 NumTotalTilesInHeap = LastHeap->GetHeapDesc().SizeInBytes / TileSizeInBytes;
        const uint32 NumUsedTilesInHeap  = NumTotalTilesInHeap - NumSlackTiles;
        const uint32 HeapFirstTile       = NumCommittedTiles - NumUsedTilesInHeap;   // 这个 heap 从哪个资源 tile 开始
 
        // ── 2. 本轮要解映射的区间 = [max(HeapFirstTile, 目标), 当前水位) ──
        const uint32 RegionEnd   = NumCommittedTiles;
        const uint32 RegionBegin = FMath::Max(HeapFirstTile, NumRequiredCommitTiles);
 
        D3D12_TILE_REGION_SIZE RegionSize = {};
        RegionSize.UseBox   = false;                   // 线性模式,可跨 mip/slice
        RegionSize.NumTiles = RegionEnd - RegionBegin;
 
        // ── 3. 记一条「映射到 NULL」的指令 ──
        FD3D12UpdateTileMappingsParams Params = {};
        Params.RangeFlags = D3D12_TILE_RANGE_FLAG_NULL;   // Heap 保持 nullptr
        Params.Coord      = GetTiledResourceCoordinate(RegionBegin, RegionSize.NumTiles);
        Params.Size       = RegionSize;
        MappingParams.Add(Params);
 
        // ── 4. 这个 heap 是全空了还是半空 ──
        if (HeapFirstTile == RegionBegin)          // 整个 heap 都被解掉 → 释放
        {
            DEC_MEMORY_STAT_BY(STAT_D3D12ReservedResourcePhysical, LastHeap->GetHeapDesc().SizeInBytes);
            LastHeap->DeferDelete();
            BackingHeaps.Pop();
            /* 同步弹出对应数量的 residency handle */
            NumSlackTiles = 0;                     // 新的 Last 是满的
        }
        else                                       // 只解了尾巴 → 留着,记 slack
        {
            NumSlackTiles += RegionSize.NumTiles;
        }
 
        NumCommittedTiles -= RegionSize.NumTiles;
    }
}

为什么是 while 循环? 因为一次收缩可能跨越多个 heap(比如从 640 掉到 300,要先整个干掉 Heap C,再切掉 Heap B 的一部分)。每轮只处理「当前最后一个 heap」。

HeapFirstTile 那行是全函数最容易看晕的:

NumCommittedTiles = 640, LastHeap = C(128 tiles), NumSlackTiles = 0
NumUsedTilesInHeap = 128 - 0 = 128
HeapFirstTile      = 640 - 128 = 512     ← Heap C 服务的是资源 tile [512, 640)

本质就是「从水位线往回倒推最后一个 heap 的起点」。

阶段 5B:增长分支(L454-551)

else   // 要变大(或不变,循环直接不进)
{
    while (NumCommittedTiles < NumRequiredCommitTiles)
    {
        const uint32 NumRemainingTiles = NumRequiredCommitTiles - NumCommittedTiles;
        ID3D12Heap* D3DHeap = nullptr;
        uint32 HeapRangeStartOffsetInTiles = 0;
        D3D12_TILE_REGION_SIZE RegionSize = {};
        RegionSize.UseBox = false;
 
        if (NumSlackTiles)                      // ── 路径 A:先吃掉现有 heap 的余量 ──
        {
            const auto& LastHeap = BackingHeaps.Last();
            const uint32 NumTotalTilesInHeap = LastHeap->GetHeapDesc().SizeInBytes / TileSizeInBytes;
 
            RegionSize.NumTiles         = FMath::Min(NumSlackTiles, NumRemainingTiles);
            HeapRangeStartOffsetInTiles = NumTotalTilesInHeap - NumSlackTiles;   // ★ 从 heap 里第几个 tile 开始
            D3DHeap                     = LastHeap->GetHeap();
 
            NumSlackTiles -= RegionSize.NumTiles;
            UsedResidencyHandles.Append(LastHeap->GetResidencyHandles());
        }
        else                                    // ── 路径 B:新建一个 heap ──
        {
            RegionSize.NumTiles         = FMath::Min(MaxTilesPerHeap, NumRemainingTiles);  // 一次最多 16MB
            HeapRangeStartOffsetInTiles = 0;
 
            NewHeapDesc.SizeInBytes = RegionSize.NumTiles * TileSizeInBytes;   // 精确尺寸,无 slack
            VERIFYD3D12RESULT(D3DDevice->CreateHeap(&NewHeapDesc, IID_PPV_ARGS(&D3DHeap)));
            INC_MEMORY_STAT_BY(STAT_D3D12ReservedResourcePhysical, NewHeapDesc.SizeInBytes);
            /* 包装成 FD3D12Heap、BeginTrackingResidency、追加到 BackingHeaps 和 handle 表 */
        }
 
        FD3D12UpdateTileMappingsParams Params = {};
        Params.RangeFlags        = D3D12_TILE_RANGE_FLAG_NONE;   // 顺序映射
        Params.Coord             = GetTiledResourceCoordinate(NumCommittedTiles, RegionSize.NumTiles);
        Params.Size              = RegionSize;
        Params.Heap              = D3DHeap;
        Params.HeapOffsetInTiles = HeapRangeStartOffsetInTiles;
        MappingParams.Add(Params);
 
        NumCommittedTiles += RegionSize.NumTiles;
    }
}

HeapRangeStartOffsetInTiles = NumTotalTilesInHeap - NumSlackTiles 为什么对?§4.6 会用具体数字验算。

阶段 6:真正调 D3D(L553-610)

// ① 先让所有用到的 heap 驻留
if (GEnableResidencyManagement && !UsedResidencyHandles.IsEmpty())
{
    FD3D12ResidencySet* ResidencySet = ResidencyManager.CreateResidencySet();
    ResidencySet->Open();
    for (FD3D12ResidencyHandle* Handle : UsedResidencyHandles) ResidencySet->Insert(Handle);
    ResidencySet->Close();
    ResidencyManager.MakeResident(D3DCommandQueue, MoveTemp(ResidencySet));
}
 
// ② 再逐条改页表
for (const FD3D12UpdateTileMappingsParams& Params : MappingParams)
{
    D3DCommandQueue->UpdateTileMappings(GetResource(),
        1 /*NumRegions*/, &Params.Coord, &Params.Size,
        Params.Heap,
        1 /*NumRanges*/, &Params.RangeFlags, &Params.HeapOffsetInTiles, &Params.Size.NumTiles,
        D3D12_TILE_MAPPING_FLAG_NONE);
}
 
// ③ 告诉 residency manager「这批 heap 已被队列引用」
// Signal the fence for this queue after UpdateTileMappings complete.
// This is analogous to executing a command list that references a set of resources.
ResidencyManager.SignalFence(D3DCommandQueue);

三步的顺序不能换:因为 reserved resource 本身不做 residency 追踪(见 §3.2 的 NOTE),驻留单位是 backing heap,必须先 MakeResident 才能映射过去。

这里必须分清两个不同的问题(对应 §2.1 分层图的 tile mapping 层与 residency 层):

回答的问题由谁负责
Tile MappingResource virtual tile → 哪个 Heap 的哪个 tile?ID3D12CommandQueue::UpdateTileMappings()
Residency提供 backing 的这个 Heap,当前在不在 GPU 可访问的驻留集合里?UE 的 Residency Manager

正确时序只有一条:创建 Heap → 让 Heap resident → 建立 tile 映射 → GPU 才能安全访问新区域。最后那次 SignalFence 的作用,是告诉 Residency Manager「这批 heap 已被这个 queue 引用到了这个时间点」,避免它过早把 heap 移出驻留集合或回收。

这个 10 参数的签名,按「左右两侧」读就清楚了:

参数含义
左侧:Resource 虚拟 tile 区间pResource / pResourceRegionStartCoordinates / pResourceRegionSizes从资源的 Coord 开始,连续取 N 个虚拟 tile
右侧:Heap tile 区间pHeap / pRangeFlags / pHeapRangeStartOffsets / pRangeTileCounts从指定 Heap 的 HeapOffset 开始,连续取 N 个 backing tile

一次普通映射的精确语义,就是把两侧一一对应起来:

Map(ResourceStartTile = X, NumTiles = N, Heap = HeapK, HeapStartTile = H)

    对 k = 0 … N-1:   R[X + k] → HeapK[H + k]

而 RangeFlags = NULL 时右侧整个不存在,语义退化成 R[X + k] → NULL。它改的自始至终只是地址翻译关系,不搬运任何数据 —— 这也是「新提交区域内容未定义」的根本原因。

参数配对关系上,「区域侧的总 tile 数」必须等于「范围侧的总 tile 数」。UE 每次只发一个区域 + 一个范围,所以直接把同一个 Params.Size.NumTiles 同时当作区域大小和范围计数传进去 —— 这就是 &Params.Size.NumTiles 出现在 pRangeTileCounts 位置的原因。

⚠️ 这里其实可以把 MappingParams 打包成一次 UpdateTileMappings(NumRegions=N, NumRanges=N) 调用,减少驱动调用次数。UE 选了逐条发,应该是为了代码简单 + 单次 commit 的 params 通常只有 1~3 条(16MB 粒度已经把次数压得很低了)。

最后更新统计:

if (ReservedResourceData->NumCommittedTiles != NumPreviousCommittedTiles)
{
    int64 CommitDeltaInBytes = TileSizeInBytes * FMath::Abs((int32)NumCommittedTiles - (int32)NumPreviousCommittedTiles);
    UE::RHICore::UpdateReservedResourceStatsOnCommit(CommitDeltaInBytes, bBuffer,
                                                     NumCommittedTiles > NumPreviousCommittedTiles);
}

UpdateReservedResourceStatsOnCommit 干的事是把这批字节数从 ReservedUncommitted* 挪到 ReservedCommitted*(RHICoreStats.cpp:195)—— 一个 INC 一个 DEC,所以两个统计量之和恒等于预留总量。§9 就是靠这个差值判断省了多少显存。

4.5 GetTiledResourceCoordinate:线性序号 → D3D 坐标

这个 lambda 回答一个问题:「资源里第 K 个 tile」在 D3D 眼里是 (Subresource, X, Y, Z) 的哪一个?

D3D 不认识「第 K 个 tile」,它只认识「第几个 subresource 的第 (x,y,z) 个 tile」。而 UE 的整套水位线逻辑是线性的,所以必须有这个翻译层。

先破除一个由函数名引起的误解:它只计算 Resource 侧的坐标,既不返回 Heap,也不返回 heap offset。 映射的两侧是由三个不同参数分别描述的:

参数属于哪个坐标空间
D3D12_TILED_RESOURCE_COORDINATE CoordResource 虚拟 tile 坐标 —— 这个 lambda 的唯一产物
HeapOffsetInTilesHeap 内的 backing tile 起点 —— 由阶段 5 的两条路径决定
NumTiles两侧连续对应的数量

前提:tile 的线性排布规则

资源整体(假设 2 个 array slice)
┌──────────────────── slice 0 ────────────────────┬──────── slice 1 ────────┐
│ mip0 tiles │ mip1 tiles │ mip2 tiles │ mip tail │ mip0 │ mip1 │ ... │tail │
└─────────── NumTotalTilesPerArraySlice ──────────┴─────────────────────────┘

所以:

const uint32 ArraySliceIndex       = OffsetInTiles / NumTotalTilesPerArraySlice;
const uint32 TileIndexInArraySlice = OffsetInTiles % NumTotalTilesPerArraySlice;

第一步:确定落在哪一级 mip

uint32 MipLevel = 0;
uint32 NextMipTileThreshold = 0;
while (MipLevel < PackedMipDesc.NumStandardMips)
{
    const D3D12_SUBRESOURCE_TILING& CurrentMipTiling = MipTilingInfo[MipLevel];
    NextMipTileThreshold += CurrentMipTiling.WidthInTiles * CurrentMipTiling.HeightInTiles * CurrentMipTiling.DepthInTiles;
    if (TileIndexInArraySlice < NextMipTileThreshold) break;   // 落在本级
    MipLevel += 1;
}

如果一路没 break,循环自然停在 MipLevel == NumStandardMips —— 这正好就是「第一个 packed mip」的索引,下面靠这个区分两条路。

ResourceCoordinate.Subresource = MipLevel + ArraySliceIndex * NumTotalMips;

D3D12 的 subresource 编号规则就是 MipSlice + ArraySlice * MipLevels。

第二步 · 情况 A:标准 mip

const uint32 NumTilesPerVolumeSlice = CurrentMipTiling.WidthInTiles * CurrentMipTiling.HeightInTiles;
const uint32 TileIndexInMipLevel    = TileIndexInArraySlice - CurrentMipTiling.StartTileIndexInOverallResource;
 
ResourceCoordinate.X = TileIndexInMipLevel % CurrentMipTiling.WidthInTiles;
ResourceCoordinate.Y = (TileIndexInMipLevel / CurrentMipTiling.WidthInTiles) % CurrentMipTiling.HeightInTiles;
ResourceCoordinate.Z = TileIndexInMipLevel / NumTilesPerVolumeSlice;

就是标准的一维序号 → 三维下标解包(行主序)。StartTileIndexInOverallResource 是这级 mip 在 slice 内的起始 tile 号,减掉它就得到 mip 内的局部序号。

第二步 · 情况 B:packed mip(mip tail)

checkf(NumTiles <= MaxTilesPerHeap,
       TEXT("Reserved texture packed mip level requires tiles: %d, maximum supported tiles: %d. ")
       TEXT("Increase d3d12.ReservedResourceHeapSizeMB or avoid packed mips by using a larger texture dimensions."),
       NumTiles, MaxTilesPerHeap);
 
// Entire packed mip chain must be covered in one map operation, so mapping origin is always 0
ResourceCoordinate.X = 0;
ResourceCoordinate.Y = 0;
ResourceCoordinate.Z = 0;

mip tail 没有几何形状可言(驱动内部怎么塞的是黑箱),只能整体映射。官方规定是:当区域包含非标准 tiling 的 mip 时 UseBox 必须为 FALSE,起始位置用 x 在扁平 tile 范围里偏移、y 和 z 必须为 0。UE 因为总是整条一起映射,所以 x 也固定为 0。

那句 assert 的实际含义:mip tail 必须能被一个 heap 一次性覆盖(因为一次 UpdateTileMappings 只能指向一个 heap)。这正是 §2.2 那条「packed mip 必须一次映射完」的直接后果。如果 mip tail 超过 16MB,就会撞上这个 check —— 报错信息也直接给了两条出路:调大 d3d12.ReservedResourceHeapSizeMB,或者把纹理做大到不产生 packed mip。

顺带解释这个 lambda 的第二个参数 NumTiles:它在 buffer 与标准 mip 路径下完全不影响返回的坐标,唯一的用处就是上面这句 packed-mip 的约束检查。第一次读的时候可以当它不存在。

4.6 数字走查(一):Buffer 的增长 → 收缩 → 再增长

设定:2GB reserved buffer,d3d12.ReservedResourceHeapSizeMB = 16

TileSizeInBytes     = 65536 (64KB)
D3DResourceNumTiles = 2GB / 64KB  = 32768
MaxTilesPerHeap     = 16MB / 64KB = 256
bBuffer             = true → NumStandardMips 强制为 1
MipTilingInfo[0]    = { WidthInTiles=32768, Height=1, Depth=1, Start=0 }
NumTotalTilesPerArraySlice = 32768

① 首次提交 40MB(0 → 640 tiles)

NumRequiredCommitTiles = 40MB / 64KB = 640

轮次slack动作RegionSizeCoord (X)Heap / offset之后 NumCommittedTiles
10建 Heap A (16MB)256GTRC(0) → X=0A / 0256
20建 Heap B (16MB)256GTRC(256) → X=256B / 0512
30建 Heap C (8MB)128GTRC(512) → X=512C / 0640

第三个 heap 只建了 8MB(128 × 64KB),不是 16MB —— 所以此时 slack 仍为 0。

坐标怎么算的(以 GTRC(256) 为例):ArraySliceIndex = 256/32768 = 0;mip 循环第一轮 NextMipTileThreshold = 32768,256 < 32768 成立直接 break → MipLevel = 0;Subresource = 0;X = (256 - 0) % 32768 = 256,Y = 0,Z = 0。

此时的概念页表:

R[0]   → Heap A[0]        R[256] → Heap B[0]        R[512] → Heap C[0]
R[1]   → Heap A[1]        R[257] → Heap B[1]        R[513] → Heap C[1]
...                       ...                       ...
R[255] → Heap A[255]      R[511] → Heap B[255]      R[639] → Heap C[127]

R[640] … R[32767]  → NULL / unmapped

NumCommittedTiles = 640,   NumSlackTiles = 0

①′ 反过来看:一次读访问是怎么翻译过去的

上面全是「怎么建立映射」。反过来,shader 读这个 buffer 的第 17MB 处会发生什么:

ByteOffset = 17MB
TileIndex  = 17MB / 64KB = 272

页表查询:  R[272] → Heap B[272 - 256] = Heap B[16]
Heap 内偏移:16 × 64KB = 1MB

完整翻译路径:
  ReservedBuffer GPUVA + 17MB
        ↓
  Resource virtual tile R[272]
        ↓  tile mapping
  Heap B 的第 16 个 tile
        ↓
  Heap B 内偏移 1MB

Shader 只看到一个连续的 buffer,完全不知道 17MB 处已经跨到了第二个 heap。 这就是 reserved resource 全部收益的来源:物理侧可以碎成 N 段、可以随时增删,虚拟侧的地址和 descriptor 岿然不动。

② 收缩到 300 tiles(640 → 300)

轮次LastHeapNumTotalTilesInHeapslack(前)UsedHeapFirstTileRegionBegin解映射 tile 数heap 处置slack(后)水位
1C1280128640-128=512max(512,300)=512640-512=128512 == 512 → 删除 C0512
2B2560256512-256=256max(256,300)=300512-300=212256 != 300 → 保留 B212300

结果:BackingHeaps = [A, B],NumCommittedTiles = 300,NumSlackTiles = 212。Heap B 只有前 44 个 tile 在用(对应资源 tile 256~299)。

❗ 这里必须分清两件事:

R[300..511] → NULL         ✅ 发生了:资源侧的映射被清除
Heap B[44..255] 被删除      ❌ 没有发生:这段物理内存仍属于 Heap B,只是没有资源 tile 指向它了

D3D12_TILE_RANGE_FLAG_NULL 清除的是资源 → heap 的连接,不销毁 heap 对象。只有当整个尾部 heap 都不再被任何资源 tile 指向时(本例中的 Heap C),UE 才 DeferDelete() 它。这也正是 NumSlackTiles = 212 的物理含义:已经付了钱、但暂时没人用的 13.25MB。

③ 再增长到 400 tiles(300 → 400)

NumRemainingTiles = 100,NumSlackTiles = 212 > 0  → 走「路径 A:吃 slack」

RegionSize.NumTiles         = min(212, 100) = 100
HeapRangeStartOffsetInTiles = 256 - 212     = 44
Coord                       = GTRC(300)     → X = 300
NumSlackTiles               = 212 - 100     = 112
NumCommittedTiles           = 300 + 100     = 400

验算这个 44 对不对:Heap B 服务的是资源 tile [256, 512)。资源 tile 300 应该落在 Heap B 的第 300 - 256 = 44 个 tile。完全吻合,一次 CreateHeap 都没多花。

4.7 数字走查(二):纹理的坐标换算

设定:512×512、RGBA8(32bpp)、10 级完整 mip、非数组。32bpp 的标准 tile 是 128×128 texel。

mip0  512×512 → 4×4 = 16 tiles,  Start = 0
mip1  256×256 → 2×2 =  4 tiles,  Start = 16
mip2  128×128 → 1×1 =  1 tile,   Start = 20
mip3~9 (64×64 及以下) → 小于一个 tile → 打包成 mip tail,占 1 tile

NumStandardMips = 3, NumPackedMips = 7, NumTilesForPackedMips = 1
D3DResourceNumTiles = 16 + 4 + 1 + 1 = 22
NumTotalTilesPerArraySlice = 21 + 1 = 22
NumTotalMips = MipTilingInfo.Num() = 3 + 7 = 10

GetTiledResourceCoordinate(18)

ArraySliceIndex       = 18 / 22 = 0
TileIndexInArraySlice = 18 % 22 = 18

mip 循环:
  MipLevel=0: threshold += 16 → 16;  18 < 16 ? 否 → MipLevel=1
  MipLevel=1: threshold +=  4 → 20;  18 < 20 ? 是 → break
→ MipLevel = 1,Subresource = 1 + 0×10 = 1

标准 mip 路径(1 < 3):
  T = mip1 { Width=2, Height=2, Depth=1, Start=16 }
  TileIndexInMipLevel = 18 - 16 = 2
  X = 2 % 2 = 0
  Y = (2 / 2) % 2 = 1
  Z = 2 / (2×2) = 0
→ (Subresource=1, X=0, Y=1, Z=0)

翻译成人话:mip1 里第 2 行第 1 列的那个 128×128 区块。mip1 是 256×256,正好 2×2 个 tile,index 2 就是左下角。

GetTiledResourceCoordinate(21)

TileIndexInArraySlice = 21
mip 循环:
  MipLevel=0: threshold=16;  21<16 否 → 1
  MipLevel=1: threshold=20;  21<20 否 → 2
  MipLevel=2: threshold=21;  21<21 否 → 3
  MipLevel(3) < NumStandardMips(3) ? 否 → 循环退出
→ MipLevel = 3 = NumStandardMips  → 走 packed 分支
→ (Subresource=3, X=0, Y=0, Z=0),并断言这次 NumTiles ≤ MaxTilesPerHeap

Subresource 3 就是第一个 packed mip,坐标 (0,0,0) + 线性 NumTiles 覆盖整条 mip tail。

五、Buffer 与 Texture 的区分

上一节的水位线模型对 buffer 完美适用。但对纹理,UE 只支持创建时一次性全量提交(TexCreate_ImmediateCommit),运行时不能 commit / decommit。这不是文档里的一句话,是四层代码同时堵死的。

5.1 四条硬证据

证据 1:RHI 验证层直接把纹理堵死了

// Engine/Source/Runtime/RHI/Private/RHIValidation.cpp:930-944
if (const FRHICommitResourceInfo* CommitInfo = Info.CommitInfo.GetPtrOrNull())
{
    if (Info.Type == FRHITransitionInfo::EType::Buffer)
    {
        RHI_VALIDATION_CHECK(EnumHasAllFlags(BufferUsage, BUF_ReservedResource), ...);
        RHI_VALIDATION_CHECK(CommitInfo->SizeInBytes <= BufferSize, ...);
    }
    else
    {
        RHI_VALIDATION_CHECK(false, TEXT("Reserved resource commit is only supported for buffers"));
    }
}

证据 2:D3D12 后端连分支都没写

// Engine/Source/Runtime/D3D12RHI/Private/D3D12LegacyBarriers.cpp:1185
// Enhanced Barriers 版本在 D3D12EnhancedBarriers.cpp:3303 是逐字相同的一份
if (Info.Type == FRHITransitionInfo::EType::Buffer)
{
    FD3D12Buffer* Buffer = Context.RetrieveObject<FD3D12Buffer>(Info.Buffer);
    Context.SetReservedBufferCommitSize(Buffer, CommitInfo->SizeInBytes);
}
else
{
    checkNoEntry();          // ← 纹理走到这里直接崩
}

而且函数名就叫 SetReserved**Buffer**CommitSize(FD3D12**Buffer*** Buffer, ...),签名层面就只收 buffer。FRHITransitionInfo 里带 FRHICommitResourceInfo 的构造函数也只有 buffer 版本(RHITransition.h:178)。

证据 3:Vulkan 后端在实现里直接断言

// Engine/Source/Runtime/VulkanRHI/Private/VulkanTexture.cpp:931
TArray<VkSparseMemoryBind> FVulkanTexture::CommitReservedResource(uint64 RequiredCommitSizeInBytes)
{
    checkf((RequiredCommitSizeInBytes == UINT64_MAX) || (RequiredCommitSizeInBytes == 0),
        TEXT("Only full resource commits are supported for images."));

证据 4:Vulkan 创建时干脆没开 residency bit

// Engine/Source/Runtime/VulkanRHI/Private/VulkanTexture.cpp:382
if (EnumHasAnyFlags(UEFlags, TexCreate_ReservedResource))
{
    ImageCreateInfo.flags |= VK_IMAGE_CREATE_SPARSE_BINDING_BIT;
    // no VK_IMAGE_CREATE_SPARSE_RESIDENCY_BIT yet since only TexCreate_ImmediateCommit is supported
}

Vulkan 里这两个 bit 分工很清楚:SPARSE_BINDING_BIT = 内存可以来自多块分配,但整张图必须全部绑定;SPARSE_RESIDENCY_BIT = 允许部分驻留(真稀疏)。UE 只开了前者。

旁证:所有纹理用例无一例外都带 ImmediateCommit

用例创建 flag
VSM Physical Page Pool(VirtualShadowMapCacheManager.cpp:1096)TexCreate_ReservedResource | TexCreate_ImmediateCommit
SVT Tile Data Texture(SparseVolumeTextureTileDataTexture.cpp:204)TexCreate_ReservedResource | TexCreate_ImmediateCommit
MediaTextureResource只是个透传参数,注释里提了一嘴

5.2 根因 A:改分辨率会打乱整个 tile↔像素映射

假设真能把 512×512「扩」成 1024×1024:

512×5121024×1024
mip0 tile 排布4×4 = 168×8 = 64
tile #4 代表什么mip0 的 (0,1) ← 第 2 行第 1 列mip0 的 (4,0) ← 第 1 行第 5 列
mip1 起点tile 16tile 64

已有数据全部错位。而 buffer 完全不同:

buffer:  元素 i 的地址 = base + i × stride     ← 扩容不改变任何已有元素的位置
texture: tile k 的含义 = f(k, 分辨率, mip 链)  ← 改分辨率 = 换了一个函数

「尾部追加」这个操作在纹理上没有对应的语义。

5.3 根因 B(更致命):尾部伸缩的方向和 mip streaming 是反的

就算不改分辨率,回看 tile 在资源里的排列顺序:

tile:  0 ─────────────── 15 │ 16 ─ 19 │ 20 │ 21
       └──── mip0 (16) ────┘ └mip1(4)┘ └m2┘ └tail┘
       ↑ 占 73% 空间,最该被卸载       最该常驻 ↑
                                                 │
                              UE 的水位线只能从这头砍 ──┘

mip0 在最前面,mip tail 在最后面。

mip streaming 想省内存,要卸载的是 mip0(占了 73% 的 tile,且远处物体根本用不到);要保留的是小 mip(永远得有个兜底)。而 UE 的模型是 NumCommittedTiles 尾部伸缩 —— 能砍的恰恰是最该留的那些,该砍的在最前面动不了。

要支持 mip streaming,就得把标量水位线换成真正的稀疏位图 + 空闲 tile 分配器,还得处理 Tier 语义(读 NULL tile 返回 0)、shader 侧的 LOD clamp / CheckAccessFullyMapped 反馈 —— 那是另一个量级的系统。UE 没有为此付这个成本。

5.4 那 CommitReservedResource 里的 mip 逻辑给谁用?

给「分段全量提交」用的,不是给「部分驻留」用的。

关键在于 heap 上限只有 16MB(d3d12.ReservedResourceHeapSizeMB)。一个 VSM page pool 动辄 256MB512MB,全量提交也得**拆成 1632 次 UpdateTileMappings**,每次的起点是不同的线性 tile 序号 —— 而 D3D 只认 (Subresource, X, Y, Z)。

所以 GetTiledResourceCoordinate 的真实职责是:

「第 4096 个 tile 在哪级 mip 的哪个位置?」—— 为了让第 17 次 UpdateTileMappings 知道从哪儿接着铺。

用 §4.7 的 512×512 例子(22 个 tile)走一遍 ImmediateCommit(RequiredCommitSizeInBytes = UINT64_MAX):

NumRequiredCommitTiles = 22(全部)
MaxTilesPerHeap = 256 → 22 < 256,一次就够
→ 建 1 个 heap(22 × 64KB = 1408KB)
→ 1 次 UpdateTileMappings:Coord = GTRC(0) = (Subresource=0, 0,0,0),NumTiles = 22
→ UseBox=false,运行时自动从 mip0 一路铺到 mip tail

只有当资源大到超过 16MB(>256 tiles)才会分多次,那时才真正用上坐标换算。VSM 的 256MB page pool 就是这种情况。

所以那套 mip 逻辑不是「稀疏能力的残留」,是「分段提交的必需品」。

5.5 纹理用 reserved 图什么:三个和稀疏无关的动机

SVT 的 CVar 描述把话说全了(SparseVolumeTextureTileDataTexture.cpp:14):

“Allocate the SVT tile data texture (streaming pool) as a reserved/virtual texture, backed by N small physical memory allocations to reduce fragmentation. This lifts the 2GB resource size limit and also allows for better GPU memory management when allocating the texture.”

#动机说明
1抗碎片256MB 一整块连续物理显存 → 16 个 16MB 小块。VSM 的注释:“This helps Windows video memory manager page allocations in and out of local memory more efficiently.”
2突破单资源尺寸上限SVT 平时受 MaxResourceSize = 2048MB 约束(SparseVolumeTextureUtility.h:15);开了 reserved 后上限改成 2048³ × 每体素字节数,只受 MaxVolumeTextureDim = 2048 这个维度限制(SparseVolumeTextureTileDataTexture.cpp:38-56)
3换页效率物理内存分散成小块,OS 可以按块换入换出,而不是整块 256MB 一起搬

注意动机 2 的连带效应:它还带了个 r.SparseVolumeTexture.Streaming.ReservedResourcesMemoryLimit(默认 -1 = 不限)作为刹车,注释说 “Without this limit it is theoretically possible to allocate enormous amounts of memory.” —— 因为一旦解开 2GB 枷锁,就得自己防着分配爆炸。

❗ SVT 这条路径的门槛比其余用例都高一档:tile data texture 是 3D volume texture,FTileDataTexture::ShouldUseReservedResources() 查的是 GRHIGlobals.ReservedResources.SupportsVolumeTextures,也就是要求 Tiled Resources Tier 3,而不是全局那个 Tier 2 的 Supported。

5.6 真要改纹理大小:销毁重建

VSM 就是这么干的(VirtualShadowMapCacheManager.cpp:1100-1109):

if (!PhysicalPagePool
    || PhysicalPagePool->GetDesc().Extent    != RequestedSize
    || PhysicalPagePool->GetDesc().ArraySize != RequestedArraySize
    || RequestedMaxPhysicalPages             != MaxPhysicalPages
    || PhysicalPagePoolCreateFlags           != RequestedCreateFlags)
{
    if (PhysicalPagePool)
    {
        UE_LOGF(LogRenderer, Display,
            "Recreating Shadow.Virtual.PhysicalPagePool due to size or flags change. This will also drop any cached pages.");
    }
    ...
}

整个池销毁重建 + 丢弃所有缓存页。 这也正好解释了为什么各家优化指南都说不要按关卡 / 按流送单元去改 r.Shadow.Virtual.MaxPhysicalPages —— 池重建不是免费的,reserved resource 一点也没让它变便宜(它优化的是分配方式,不是重建这件事本身)。

5.7 UE 里真正的「纹理稀疏」是软件虚拟化

系统实现方式
Virtual Texture (VT/RVT)自己维护 page table 纹理 + physical pool 纹理,shader 里做两次采样间接寻址
Virtual Shadow Map同上:16K×16K 是概念上的虚拟分辨率,实际存储是固定大小的 physical page pool + page table
Sparse Volume Texture同上:page table volume + physical tile data volume

三者都是把「稀疏」做在 shader 层,硬件 tiled resources 只被拿来当「抗碎片的分配器」。

⚠️ 原因没有写在任何注释里,但从代码分布看非常明显:

  • 跨平台 —— 软件方案在 Metal / 移动端 / 主机上一致可用,硬件 tiled resources 的 Tier 支持参差不齐(UE 自己就把门槛卡在 Tier 2,3D 还要 Tier 3);
  • 可控 —— page 大小、替换策略、反馈机制全在自己手里,不依赖驱动语义;
  • 调试友好 —— page table 是一张普通纹理,RenderDoc 能直接看;而 tiled resources 的抓帧支持历史上一直有缺口。

六、UE5 里的三个消费者

三个消费者在 Windows 构建上全部处于开启状态,但动机各不相同:

系统开关默认状态资源类型
GPUScener.GPUScene.UseReservedResourcesC++ 默认 trueBuffer(可伸缩)
Nanite Streamingr.Nanite.Streaming.ReservedResourcesBaseWindowsEngine.ini = 1(C++ 默认是 0)Buffer(可伸缩)
VSM Page Poolr.Shadow.Virtual.AllocatePagePoolAsReservedResourceC++ 默认 1Texture(全量提交)

6.1 GPUScene:消灭 resize 时的 double buffer + memcpy

// Engine/Source/Runtime/Renderer/Private/GPUScene.cpp:151
static TAutoConsoleVariable<bool> CVarGPUSceneUseReservedResources(
    TEXT("r.GPUScene.UseReservedResources"), true,
    TEXT("Turning this on makes all GPU-Scene buffers default to using reserved resource allocations."), ECVF_ReadOnly);

做法是一次性预留 2GB VA,之后只 commit 实际需要的部分(GPUScene.cpp:952):

constexpr uint64 ReservedAllocationSize = uint64(2048) << 20;   // 2GB of address space
BufferDescReserved.Usage |= EBufferUsageFlags::ReservedResource;

收益体现在这个分岔上(UnifiedBuffer.cpp:505):

if (EnumHasAllFlags(ExternalBuffer->Desc.Usage, EBufferUsageFlags::ReservedResource) && ...)
{
    GraphBuilder.QueueCommitReservedBuffer(InternalBufferOld, BufferSizeNew);
    return InternalBufferOld;          // ← 同一个 buffer、同一个 GPU VA、descriptor 全部不失效
}
else
{
    InternalBufferNew = GraphBuilder.CreateBuffer(...);                          // 新建
    MemcpyResource(GraphBuilder, InternalBufferNew, InternalBufferOld, Params);  // + 全量拷贝
    return InternalBufferNew;          // ← 峰值显存 = 旧 + 新,且所有 view 要重建
}

Epic 在 5.8 的说明也直白确认了这个动机:“Enabled Reserved Resources for GPU Scene to reduce hitches on resize and peak GPU memory use by eliminating the need to double buffer during the copy.”

6.2 Nanite Streaming Pool:同样是「可增长池」

// Engine/Source/Runtime/Engine/Private/Rendering/NaniteStreamingManager.cpp:188
static int32 GNaniteStreamingReservedResources = 0;   // ← C++ 默认值是 0
static FAutoConsoleVariableRef CVarNaniteStreamingReservedResources(
    TEXT("r.Nanite.Streaming.ReservedResources"),
    GNaniteStreamingReservedResources,
    TEXT("Allow allocating Nanite GPU resources as reserved resources for better memory utilization and more efficient resizing (EXPERIMENTAL)"),
    ECVF_ReadOnly | ECVF_RenderThreadSafe
);

❗ 只看 C++ 默认值会得出错误结论。[SystemSettings] 会在启动时覆盖它:

; Engine/Config/Windows/BaseWindowsEngine.ini:23,在 [SystemSettings] 段内
r.Nanite.Streaming.ReservedResources=1

这是上游 Epic 的平台默认设置,所以 Windows 上 Nanite 的 reserved resource 路径是开着的。

这条 ini 精确做了什么

维度结论
生效范围Windows 平台的 Editor + Game 都吃;项目 DefaultEngine.ini 的 [SystemSettings] 可覆盖
可否运行时改❌ CVar 标记为 ECVF_ReadOnly | ECVF_RenderThreadSafe。控制台改不动,只能改 ini 或启动加 -dpcvars=r.Nanite.Streaming.ReservedResources=0
只是「允许」,不是「强制」实际生效还要 GRHIGlobals.ReservedResources.Supported:bReservedResource = (Supported && GNaniteStreamingReservedResources)(NaniteStreamingManager.cpp:550)
不区分 RHID3D12 → Tiled Tier ≥ 2 才 true;Vulkan on Windows 也会命中(需 sparseBinding + sparseResidencyBuffer + sparseResidencyImage2D + residencyNonResidentStrict,且 graphics queue 有 VK_QUEUE_SPARSE_BINDING_BIT,且 SupportsParallelRendering());D3D11 下静默 fallback
作用对象只有 Nanite.StreamingManager.ClusterPageData 这一个 buffer。Hierarchy.DataBuffer 仍是普通 buffer(NaniteStreamingManager.cpp:573)

⚠️ CVar 描述里写着 EXPERIMENTAL,但 Epic 已经在平台 ini 里默认打开了;标签看起来是历史遗留没清理,实际成熟度可参照 GPUScene(5.8 已在 release notes 里正式宣传)。

开启后 Nanite 内部行为的真实差异

对应 NaniteStreamingManager.cpp:1673-1833:

关闭(纯 C++ 默认值)开启(Windows 现状)
buffer 创建CreateByteAddressDesc(4),靠 resize 长大一次性预留 GetMaxPagePoolSizeInMB() 的 VA
root page 分配器FSpanAllocator(true) 强制 grow-only允许收缩(除非 r.Nanite.Streaming.Debug.ReservedResourceRootPageGrowOnly=1)
初始 root pageNumInitialRootPages = 2048 → 2048×32KB = 64MB 起步从 0 起步(ReservedResourceIgnoreInitialRootAllocation=1 默认开)
增长粒度RoundUpToSignificantBits(N, 2),跳跃式固定 16MB chunk(// Allocate pages in 16MB chunks to reduce the number of page table updates)
池 resize 的实现新建 buffer + AddCopyBufferPass 搬 root page,两份 allocation 同时存活原地 AddPass_Memmove + commit/decommit,无临时峰值
CSV 事件发 GrowPoolAllocation不发(因为没有真正的 grow)

Epic 自己在非 reserved 分支留的注释很说明问题:

// Non-reserved resource path: Make new allocation and copy root pages over.
// Temporary peak in memory usage when both allocations need to be live at the same time.
// ...
//   It might not be worthwhile if reserved resources will be supported on all relevant platforms soon.

而 reserved 分支要注意 memmove 的方向依赖 resize 方向(缩小时先搬后 resize,放大时先 resize 后搬)—— 因为 root page 就住在 streaming page 区域的上方,同一个 buffer 里,commit 是尾部伸缩的,顺序错了会踩到未提交的 tile。

一个调参上的连带变化:r.Nanite.Streaming.StreamingPoolSize(默认 512MB)现在改它不再触发 buffer 重建 + 全量拷贝,只是 memmove + 改页表。但仍然会重置流送状态 —— MaxStreamingPages 变化会走 bResetStreamingState 分支,UninstallAllResidentPages() 全部卸载并 memset 清零。所以「运行时调 pool size 会掉细节、要重新流送」这个现象依然存在,只是不再有显存尖峰和 hitch。

6.3 VSM Physical Page Pool(与 SVT):动机完全不同

// Engine/Source/Runtime/Renderer/Private/VirtualShadowMaps/VirtualShadowMapCacheManager.cpp:1094
// Using ReservedResource|ImmediateCommit flags hint to the RHI that the resource can be allocated using N small
// physical memory allocations, instead of a single large contighous allocation. This helps Windows video memory
// manager page allocations in and out of local memory more efficiently.
ETextureCreateFlags RequestedCreateFlags = (CVarVSMReservedResource.GetValueOnRenderThread() && GRHIGlobals.ReservedResources.Supported)
    ? (TexCreate_ReservedResource | TexCreate_ImmediateCommit) : TexCreate_None;

CVar r.Shadow.Virtual.AllocatePagePoolAsReservedResource 默认 1。

注意它带了 ImmediateCommit —— 创建时就全量提交(D3D12Texture.cpp:538):

if (EnumHasAllFlags(Flags, TexCreate_ImmediateCommit))
{
    // NOTE: Accessing the queue from this thread is OK, as D3D12 runtime acquires a lock around all command queue APIs.
    Resource->CommitReservedResource(pDevice->GetQueue(ED3D12QueueType::Direct).D3DCommandQueue, UINT64_MAX /*commit entire resource*/);
}

所以这里根本没用到「稀疏」这个能力,纯粹是借 reserved resource 把一块几百 MB 的大分配拆成 N 个 16MB heap,让 Windows 显存管理器换页更顺、避免大块连续显存申请失败。这是一个很值得学的「非典型用法」,动机见 §5.5。

同类还有 SparseVolumeTexture 的 tile data texture。有学术论文专门利用这点在 UE 里做超大体数据:“the backing memory can actually be allocated as a collection of many tiles… allowing a huge texture allocation even with fragmented VRAM”。


七、VA 账单与调优旋钮

7.1 一台 PC 实际预留了多少地址空间

Nanite 侧(NaniteStreamingManager.cpp:299):

static uint32 GetMaxPagePoolSizeInMB()
{
    const uint32 DesiredSizeInMB = IsRHIDeviceAMD() ? 4095 : 2048;
    const uint32 MaxSizeInMB = (uint32)(GRHIGlobals.MaxViewSizeBytesForNonTypedBuffer >> 20);
    return FMath::Min(DesiredSizeInMB, MaxSizeInMB);
}

D3D12RHI 没有覆写 MaxViewSizeBytesForNonTypedBuffer,走 RHI 默认 1ULL << 32 = 4096MB(RHIGlobals.h:324)。所以:

  • NVIDIA / Intel → Nanite 预留 2048 MB VA
  • AMD → 预留 4095 MB VA

再叠加 GPUScene(r.GPUScene.UseReservedResources C++ 默认就是 true,无需 ini),每个 reserved buffer 固定预留 2GB:

BufferVA备注
GPUScene.PrimitiveData2 GB
GPUScene.InstanceSceneData2 GB仅 tiled 布局时(r.GPUScene.InstanceDataTileSizeLog2 默认 12;设为负数即关闭 tiled 布局并连带禁用 reserved)
GPUScene.InstancePayloadData2 GB
GPUScene.LightmapData2 GB
Nanite.StreamingManager.ClusterPageData2 GB / 4 GB

GPUScene.LightData 不在此列 —— 它走的是普通 buffer 路径。

合计约 10 GB(N 卡)~ 12 GB(A 卡)的 GPU 虚拟地址空间,物理显存占用是另一回事(只算 commit 的部分)。

对照告警阈值 rhi.ReservedResources.VirtualSizeWarningGB 默认 256GB(RHICoreStats.cpp:11)—— PC 上完全安全。⚠️ 这个默认值明显是给主机 / 受限 VA 平台留的余量,PC 上基本触发不了;真要触发说明有代码在循环创建 reserved 资源没释放,那时它是个很好的泄漏探针。

7.2 「16MB」这个数字为什么到处都是

GPUScene 的 tiled instance data 用了和 Nanite 一模一样的 16MB 对齐(GPUScene.cpp:988):

// Grow/shrink in 16MB chunks to reduce the number of page table updates or resizes
const uint32 AllocationStrideInBytes = 16u << 20;

和 d3d12.ReservedResourceHeapSizeMB = 16 正好对齐 —— 一次 commit 恰好对应一个新 backing heap,不产生跨 heap 的碎片映射。这不是巧合。

背后的原则是 §2.5 那条:页表更新本身有成本,要按较粗粒度批量做。而且在 UE 的实现里还多两层代价 —— commit 会 CloseCommandList(),提交线程还要为它单独 Flush() 一次(§3.3)。Nanite + GPUScene 同帧各自 commit 一次,就会把该帧的提交批次切碎。这正是两边都用 16MB 粗粒度、且 Nanite 用 IgnoreInitialRootAllocation 从 0 起步按需长的原因:commit 次数比 commit 大小更值得优化。

7.3 backing heap 的对象数量

2048MB 池全提交时 = 128 个 16MB heap(Nanite 一家),加上 GPUScene 四个 buffer。这些 heap 都进 residency 管理的 handle 列表。

⚠️ 如果在 Insights 里看到 CommitReservedResource / residency MakeResident 的 CPU 时间偏高,d3d12.ReservedResourceHeapSizeMB 调到 32 或 64 是第一个该试的旋钮(ECVF_ReadOnly,需启动时设)。Vulkan 后端对应 r.Vulkan.SparseImageAllocSizeMB。


八、收益与代价清单

收益

  1. Resize 零拷贝:grow/shrink 只改页表,数据原地不动,GPU VA 不变 → 所有 SRV/UAV/descriptor 不失效(bindless 场景尤其值钱)。
  2. 削峰:省掉 resize 瞬间「旧 + 新」双份显存。
  3. 抗碎片:大资源由 N 个小 heap 拼成,不要求一块连续物理显存。
  4. 突破单资源尺寸上限:SVT 借此绕开 2GB MaxResourceSize(§5.5)。
  5. VA 免费:预留 2GB 地址空间不花任何物理显存。
  6. 真稀疏(UE 目前基本没用):VT / SVT / 超大体纹理的天然载体。

代价与坑

  1. VA 是有预算的。UE 专门加了告警 rhi.ReservedResources.VirtualSizeWarningGB:“Some platforms have a limited virtual address space for reserved resources”。GPUScene 一个 buffer 就吃 2GB VA,多开几个要留意(§7.1)。
  2. ❗ 64KB 粒度取整:AlignArbitrary(RequiredCommitSizeInBytes, 65536) —— 请求 1 字节实际占 1 个 tile(64KB),请求 65KB 实际占 2 个 tile(128KB)。小资源用它是净亏。
  3. packed mip 必须整体映射,且 array + 小 mip 在 Tier < 4 上不合法 → UE 用 TextureArrayMinimumMipDimension = 256 保守规避。
  4. 页表更新有真实 CPU / 驱动开销,必须批量(Nanite 的 16MB 粒度就是这个原因);并且它会切断 command list,频繁 commit 会打碎提交批次。
  5. 部分队列不支持 → 触发跨队列 fence 串行化(§3.4)。
  6. 新提交区域内容未定义,必须先写再读。
  7. 纹理不能运行时伸缩,改尺寸只能销毁重建、丢掉全部缓存(§5.6)。
  8. 压缩(DCC)可能受影响,NVIDIA 建议按 2MB 对齐更新;UE 默认 16MB heap 天然满足,但单次 commit 粒度不一定。
  9. 工具链:RenderDoc 对 D3D12 tiled resources 的支持历史上就有缺口;PIX 对 reserved 资源是按 subresource 粒度做 active/inactive 追踪的,“inactive” 的 mip 内容不会被抓下来。抓帧调试要有心理预期。

九、如何验证它在工作

① 看已提交的物理量

stat nanitestreaming

关注 Total Pool Size (MB) / Root Pool Size (MB) / Streaming Pool Size (MB) —— 这些是已提交的物理量。

② 看预留 vs 实际的差值

stat rhi

Engine/Source/Runtime/RHICore/Private/RHICoreStats.cpp 提供的口径:

  • STAT_ReservedUncommittedTextureMemory / STAT_ReservedUncommittedBufferMemory = 预留但未提交(buffer 侧应该很大,≈10GB 量级)
  • STAT_ReservedCommittedBufferMemory / STAT_ReservedCommittedTextureMemory = 已提交
  • STAT_D3D12ReservedResourcePhysical = 实际 backing heap 字节数(含 slack,所以可能略大于「已提交」)
  • GRHIGlobals.ReservedResources.VirtualSize = 所有 reserved 资源的 VA 总量

Uncommitted 与 Committed 之和恒等于预留总量(UpdateReservedResourceStatsOnCommit 是一 INC 一 DEC)。Uncommitted 这一项,就是这个特性替你省下的显存。

③ 判定是否真的走了 reserved 路径

最直接的两个办法:

  • 在 PIX / Insights 里找 CommitReservedResource 这个 TRACE_CPUPROFILER_EVENT_SCOPE;
  • 看是否还有 CSV_EVENT(NaniteStreaming, "GrowPoolAllocation") —— 走 reserved 路径时这个事件永远不会出现。

十、总结

10.1 从 UE4 心智模型迁移

UE4 你只有 「资源 = 一块内存」。DX12 把它拆成 「资源 = 一段地址 + 一张页表 + 一堆物理页」:

  • Committed = 地址和页由运行时打包给你,永久绑定;
  • Placed = 页由你管(heap),但绑定关系创建时定死、且必须连续;
  • Reserved = 连绑定关系都交给你,运行时可改、可空、可来自多个 heap、可多对一。

而 UE5 目前主要吃的不是它的「稀疏」能力,而是**「地址不变、物理内存可原地伸缩」**这一条 —— 这正好治好了 GPU-driven 渲染里最痛的那个病:大 buffer 扩容要 double buffer + 全量 memcpy + 重建所有 view。

10.2 Buffer vs Texture 一页对照

Reserved BufferReserved Texture
UE 是否支持部分提交✅ 支持,QueueCommitReservedBuffer❌ 不支持,validation 直接报错
提交时机运行时任意次,尾部伸缩仅创建时一次(ImmediateCommit 全提交)
典型使用者GPUScene 四大 buffer、Nanite ClusterPageDataVSM PagePool、SVT TileData
核心收益resize 零拷贝、descriptor 不失效、无双份峰值抗碎片、突破 2GB 单资源上限、换页效率
为什么能 / 不能伸缩一维无结构,尾部追加语义天然成立多维有结构;且 mip0 在 tile 空间最前面,尾部水位线砍不到
想「变大」怎么办直接 commit 更多 tile销毁重建(VSM 就是这么做的,会丢缓存)
硬件「真稀疏」用了吗没有(线性水位线)没有(全提交)

一句话:UE5 把 DX12 reserved resource 当成了两样东西 —— 对 buffer,它是「地址不变的可伸缩显存水位阀」;对 texture,它是「把大块分配拆成小块的抗碎片分配器」。两者都没有用到它最出名的那个能力:稀疏驻留。


十一、参考链接

官方规格

厂商 / 工具

教程 / 示例

UE5 相关