Skip to content

Latest commit

 

History

History
399 lines (296 loc) · 24.8 KB

File metadata and controls

399 lines (296 loc) · 24.8 KB

Performance Design

SnapRHI is designed for low-latency, high-throughput rendering workloads. Every abstraction is evaluated against the cost of the underlying API call it wraps. This document explains the architectural decisions that minimize overhead and how to verify performance in your own workloads.

Related: Profiling Guide | Resource Management | API Overview


Table of Contents

  1. Design Principles
  2. Per-Backend Optimizations
  3. Overhead Budget
  4. Validation Cost Model
  5. Build Configurations for Performance Work
  6. Measuring Performance

1. Design Principles

1.1 Allocation-Controlled Hot Paths

Fixed-limit command operations use stack or inline storage. Variable-size state is reserved or pooled before normal recording and retains capacity across command-buffer reuse. It grows only when a workload exceeds that retained capacity; this is deliberately documented rather than hidden behind an absolute allocation-free claim.

Backend Hot-Path Allocation Strategy
Metal Pooled ResourceBatcher instances use 256-entry inline buckets. Descriptor sets sub-allocate from a shared MTLBuffer. Large barrier lists are emitted in fixed-size chunks.
Vulkan Fixed-limit descriptor, vertex-buffer, attachment, and clear-value translation uses stack/inline arrays. Large barrier lists are emitted in fixed-size chunks. Public command pools recycle native command storage.
OpenGL A retained linear command arena placement-constructs commands contiguously. Batched copies and barriers are split into fixed-size command records.

The OpenGL command allocator enforces this at the type level — all command structs must be trivially destructible (verified by static_assert), so the entire recording buffer can be rewound without walking individual destructors. The arena can grow if CommandPoolCreateInfo::commandMemorySize is exhausted, then reuses the larger allocation on subsequent recordings.

Retained command buffers reserve common-case resource and interop-reference capacity at construction. An unusually large recording can grow those vectors; CommandPoolCreateFlags::UnretainedResources removes shared-ownership and storage work in Release builds when the application can guarantee lifetime safety itself.

1.2 Compile-Switchable Validation

Every validation check in SnapRHI is guarded by an if constexpr test against a compile-time bitmask of enabled validation tags:

#define SNAP_RHI_VALIDATE(layer, condition, reportLevel, tags, ...)                     \
    if constexpr ((static_cast<uint64_t>(tags) & snap::rhi::EnabledValidationTags) &&   \
                  (static_cast<uint64_t>(reportLevel) >= SNAP_RHI_ENABLED_REPORT_LEVEL)) { ...

When a tag is disabled (the default for Release builds), the compiler dead-code-eliminates the entire check — not branched over at runtime, but absent from the binary. Over 30 individual validation tags can be toggled independently, or all at once via SNAP_RHI_ENABLE_ALL_VALIDATION.

Release build:   record → encode → submit        (validation code is not in the binary)
Debug + validation: record → validate → encode → submit

See §4 — Validation Cost Model for details.

1.3 Aggressive Resource Pooling and Reuse

SnapRHI pools and recycles expensive objects rather than creating and destroying them per-frame:

Resource Pooling Mechanism
Fences FencePool — acquires/releases reusable internal fences across all backends
Command buffers Public CommandPool — preconstructs a bounded set of backend wrappers and recycles them across all backends
Command pools (Vulkan) Each public pool owns one native VkCommandPool; retained and individually resettable command buffers reuse native allocations
Framebuffers (OpenGL) FramebufferPool — caches GL framebuffer objects by hashed description
Descriptor memory (Metal) Sub-allocates from a single MTLBuffer with atomic block tracking
Resource batchers (Metal) Pooled ResourceBatcher instances with acquire/release lifecycle

1.4 Two Retention Modes

Applications can opt out of automatic reference retention for maximum throughput. In Release builds, CommandPoolCreateFlags::UnretainedResources skips ResourceResidencySet bookkeeping entirely. The application then guarantees every recorded resource's lifetime through GPU completion.

See Resource Management — Retention Modes for the full contract.


2. Per-Backend Optimizations

2.1 Metal

Technique Details
Descriptor sub-allocation DescriptorPool sub-allocates from a single MTLBuffer with per-block atomic tracking — no per-set MTLBuffer creation.
Resource usage batching ResourceBatcher aggregates useResource: calls with pre-sized buckets (256 initial), plus a _lastBucket cache for O(1) consecutive-add fast path.
Pooled batchers Metal Device maintains a pool of ResourceBatcher instances — zero allocation during encoder setup.

2.2 Vulkan

Technique Details
Application-sized command pools Each public CommandPool preconstructs its bounded wrapper set and owns one native VkCommandPool. Native command buffers are reused except when an unretained, non-resettable buffer returns its allocation on release.
Descriptor pool allocation Standard VkDescriptorPool block allocation — set creation is a fast pool sub-allocation.

2.3 OpenGL

Technique Details
Linear command allocator Arena-style allocator that grows by doubling (starting at 2 KB). Commands are POD structs, placement-new'd sequentially. Iterated via CommandIterator without any virtual dispatch.
Comprehensive state cache StateCache tracks all GL state (buffers, textures, samplers, programs, framebuffers, viewport, depth, stencil, blending, vertex attributes, feature toggles). Every set*() call checks the cache first via inline should*() methods — redundant gl* calls are eliminated before reaching the driver.
Framebuffer cache FramebufferPool caches GLuint framebuffer objects keyed by a hashed attachment description — identical render pass configurations reuse existing FBOs.
Command buffer pool The public CommandPool uses the common fixed-capacity recycler; each reused buffer rewinds its CPU command arena instead of recreating it.
No glGet* calls The state cache explicitly avoids glGet* queries during rendering, which would force a GPU→CPU sync and stall the pipeline. All state is tracked CPU-side.

3. Overhead Budget

This section characterizes where SnapRHI adds overhead versus a raw backend call. The analysis is split by operation category and distinguishes between backends where behavior differs materially.

3.1 Command Recording — Render & Compute Encoders

These operations are called many times per frame (inner-loop hot path).

Operation Metal Vulkan OpenGL
draw / drawIndexed Virtual call → flushes deferred dynamic-offset/resource state when dirty → native draw Virtual call → native draw Virtual call → placement-new DrawCmd into arena. Actual GL calls happen later during replay.
bindRenderPipeline Virtual call → static_cast → stores pipeline pointer + sets depth bias via Obj-C message → resourceResidencySet.track() Virtual call → static_cast → vkCmdBindPipeline → resourceResidencySet.track() Virtual call → placement-new SetRenderPipelineCmd into arena → resourceResidencySet.track()
bindDescriptorSet Virtual call → binds argument-buffer state → stores bounded dynamic offsets → retains referenced resources Virtual call → fixed stack array → vkCmdBindDescriptorSets Virtual call → placement-new BindDescriptorSetCmd + bounded dynamic-offset copy → retains referenced resources
bindVertexBuffer Virtual call → native buffer bind → track() Virtual call → one-element stack translation → vkCmdBindVertexBuffers → track() Virtual call → placement-new SetVertexBufferCmd → track()
bindVertexBuffers (multi) Virtual call → loop of native binds + track() Fixed stack arrays → vkCmdBindVertexBuffers → track() per buffer Virtual call → one bounded SetVertexBuffersCmd → track() per buffer
bindIndexBuffer Virtual call → static_cast → stores pointer + offset (no native call) → track() Virtual call → static_cast → vkCmdBindIndexBuffer → track() Virtual call → placement-new SetIndexBufferCmd → track()
setViewport Virtual call → struct conversion → [encoder setViewport:] Virtual call → static_cast → vkCmdSetViewport + vkCmdSetScissor Virtual call → placement-new SetViewportCmd. During replay: gl.viewport() through state cache shouldViewport()
setDepthBias Virtual call → stores values → conditional [encoder setDepthBias:...] Virtual call → static_cast → vkCmdSetDepthBias Virtual call → placement-new SetDepthBiasCmd
setStencilReference Virtual call → stores values → [encoder setStencilFrontReferenceValue:...] Virtual call → static_cast → vkCmdSetStencilReference Virtual call → placement-new SetStencilReferenceCmd
setBlendConstants Virtual call → [encoder setBlendColorRed:...] Virtual call → static_cast → vkCmdSetBlendConstants Virtual call → placement-new SetBlendConstantsCmd

Key observations

  • Metal and Vulkan call native APIs inline during encoding. Metal retains some dirty resource/dynamic-offset state for a bounded flush before draw or dispatch; Vulkan binds descriptor sets immediately.
  • OpenGL records all commands into a linear allocator (arena). The actual gl* calls happen during replayCommands() at submit time, where the RenderCmdPerformer walks the arena via CommandIterator (no virtual dispatch during replay — just a switch on command IDs).
  • Fixed API limits such as MaxVertexBuffers, MaxBoundDescriptorSets, and MaxAttachments size translation arrays, avoiding per-call heap allocation while making invalid ranges explicit precondition violations.

3.2 Resource Retention (track())

Every bind/record operation calls resourceResidencySet.track() per resource. In retained mode (default), this obtains a strong reference directly from the resource's weak_from_this() state and appends it to the command buffer's pre-reserved vector. It does not touch the diagnostic live-resource tracker and therefore performs no tracker mutex acquisition or hash lookup. Consecutive duplicate resources are rejected before the weak lock.

In unretained mode (CommandPoolCreateFlags::UnretainedResources), track() returns after its mode check in Release builds. In debug builds with slow validations, recording stores weak_ptr observations and submission snapshots them for fence-scoped lifetime checking instead.

Retention Mode Per-track() Cost
Retained (default) Weak-to-strong reference acquisition + pre-reserved vector::push_back; adjacent duplicates are a pointer comparison
Unretained (Release) One mode check; no reference-count or container work
Unretained (Debug + slow validations) Weak lifetime observation appended to pre-reserved diagnostic storage

Implication: Retained mode is lock-free but still pays shared ownership bookkeeping. Use CommandPoolCreateFlags::UnretainedResources only when the application can enforce lifetimes through GPU completion.

3.3 Encoder Lifecycle (Begin / End)

Phase Metal Vulkan OpenGL
beginEncoding (RenderPass) Creates MTLRenderPassDescriptor → native encoder → tracks attachments Tracks attachments → resolves inline image views/clear values → vkCmdBeginRenderPass Placement-new BeginRenderPassCmd into arena → tracks attachments
beginEncoding (Dynamic Rendering) Creates MTLRenderPassDescriptor → native encoder Tracks attachments → builds an inline bounded attachment list → vkCmdBeginRendering Placement-new BeginRenderPass1Cmd into arena
endEncoding [encoder endEncoding] → clears pipeline state vkCmdEndRenderPass or vkCmdEndRendering → resets descriptor state Placement-new EndRenderPassCmd

Vulkan attachment and clear-value translations are bounded by public API limits and remain inline. Any native image layout work remains part of the backend operation rather than a container-allocation cost.

3.4 Command Buffer Submission

Step Cost Notes
submissionTracker.tryReclaim() Mutex lock + linear scan of in-flight submissions Reclaims completed slots; bounded by max in-flight count
Texture interop processing Per-interop-texture work Only when interop textures are present
Fence acquisition (via buildFence) FencePool::acquireFence() — mutex lock + linear scan for reusable fence Pool avoids createFence on hot path
Semaphore wait/signal Per-semaphore native wait/signal Metal: Obj-C dispatch. Vulkan: native semaphore handles. OpenGL: glFenceSync
Native submit [mtlCommandBuffer commit] / vkQueueSubmit / GL command replay Core driver cost
submissionTracker.track() Retained: mutex lock + strong command-buffer/fence resolution + slot update. Slow validation: weak lifetime snapshot + slot update Retained mode, or unretained mode only when slow validation is enabled
OpenGL-specific: replayCommands() Reset GL state cache → iterate arena via CommandIterator (switch dispatch) → RenderCmdPerformer calls gl* through state cache Queue use is externally synchronized by contract; no internal queue mutex is taken. State-cache comparisons filter redundant native calls.

3.5 Resource Creation

All creation methods go through Device::createResource<>(), which:

  1. new — heap-allocates and constructs the backend implementation object, including its native API initialization.
  2. Wraps it in std::shared_ptr with a custom deleter that carries its live-resource tracker slot.
  3. Registers the object in the live-resource tracker — mutex lock + amortized vector append at a new high-water mark, or O(1) reuse of a vacant slot. Reused slots require no allocation.

The tracker stores one raw pointer and an embedded free-list link per slot. Resource deleters unregister before object destruction, allowing diagnostic snapshots to obtain weak_from_this() under the tracker mutex without storing a second weak reference per resource.

Operation Additional Backend Cost
Create texture Metal: [device newTextureWithDescriptor:]. Vulkan: vkCreateImage + vkAllocateMemory + vkBindImageMemory + image view creation. OpenGL: glGenTextures + glTexStorage*.
Create buffer Metal: [device newBufferWithLength:]. Vulkan: vkCreateBuffer + vkAllocateMemory + vkBindBufferMemory. OpenGL: glGenBuffers + glBufferData.
Create render pipeline Metal: [device newRenderPipelineStateWithDescriptor:] (may compile shaders). Vulkan: vkCreateGraphicsPipelines (may compile shaders). OpenGL: glLinkProgram + state translation.
Create sampler Metal: [device newSamplerStateWithDescriptor:]. Vulkan: vkCreateSampler. OpenGL: glGenSamplers + glSamplerParameter*.
Create descriptor set Metal: argument buffer sub-allocation from DescriptorPool. Vulkan: vkAllocateDescriptorSets from pool. OpenGL: stores binding metadata.
Create command buffer Recycles a backend wrapper from a CommandPool; normal shared_ptr/live-tracker bookkeeping still applies. Metal: fresh one-shot [commandQueue commandBuffer]. Vulkan: retained or individually resettable pools reset and reuse the same native command buffer; unretained pools without individual reset allocate after the previous native buffer was freed on release. OpenGL: reuses its reserved CommandAllocator arena.

Pipeline creation is the most expensive operation — it may trigger shader compilation. Use PipelineCache to amortize this cost.

3.6 Resource Destruction

Destruction happens when the last shared_ptr reference is released:

  1. The custom deleter unregisters its known slot — mutex lock + O(1) free-list insertion, with no lookup or allocation.
  2. (Slow validations only) Scans incomplete submissions for weak observations of the erased resource.
  3. The backend object is destroyed, or a command-buffer wrapper is released to its pool. In an unretained Vulkan pool without individual reset, release also returns the native command-buffer allocation.

3.7 Descriptor Set Updates

Descriptor writes (bindUniformBuffer, bindTexture, bindSampler, updateDescriptorSet) are virtual calls into backend-specific logic:

  • Metal: Writes into argument buffer memory (sub-allocated from DescriptorPool). No native API call — direct memory writes.
  • Vulkan: Calls vkUpdateDescriptorSets with translated write structs.
  • OpenGL: Stores binding metadata into internal arrays (applied during replay).

3.8 Buffer Map / Unmap

  • Metal: [buffer contents] + offset — returns pointer to shared memory.
  • Vulkan: vkMapMemory / vkUnmapMemory — may involve driver synchronization.
  • OpenGL: glMapBufferRange / glUnmapBuffer.

No additional SnapRHI overhead beyond the virtual call + static_cast.

3.9 Validation

Build Configuration Cost
Tag-gated validation disabled (Release default) Zero for those checks — if constexpr removes them; unrecoverable critical invariant reporting remains compiled
Validation enabled (selective tags) Per-enabled-tag checks at each guarded call site
Slow validations (SNAP_RHI_ENABLE_SLOW_VALIDATIONS) Additional per-track() weak observations + submission snapshot + lifetime scans when resources are destroyed

3.10 Summary Table

Operation Category Overhead vs. Raw API Dominant Cost
Draw / dispatch Virtual call plus native recording/replay work Metal flushes dirty bounded state; Vulkan records directly; OpenGL writes its arena
Dynamic state (viewport, depth bias, etc.) Virtual call Metal/Vulkan: inline native call. OpenGL: arena write (replay adds state cache check)
Bind vertex/index buffer Virtual call + native/recorded bind + track() Retained mode: lock-free shared ownership; fixed-limit translations remain inline
Bind descriptor set Virtual call + backend bind/record Exact dynamic-offset validation is compiled out when disabled; fixed arrays cover API limits
Begin/end render pass Virtual call + attachment setup Fixed-limit translations remain inline; native descriptor/layout work still applies
Submit command buffer Fence pool + submission tracker Submission/fence trackers synchronize shared state; OpenGL performs full command replay
Resource creation shared_ptr + live-tracker registration Mutex lock + amortized slot-vector growth + native API object creation. Pipelines: may compile shaders
Resource destruction shared_ptr custom deleter Mutex lock + O(1) slot recycle + native API release
Descriptor writes Virtual call + backend write Metal: memory write. Vulkan: vkUpdateDescriptorSets. OpenGL: store metadata

Note on virtual dispatch: SnapRHI uses virtual functions for backend polymorphism. On modern CPUs, a predicted virtual call costs ~1–2 ns. This is the cost of a uniform, type-safe multi-backend API. Applications that build against a single backend can expect the compiler to devirtualize final class methods in many cases (all backend classes are marked final).

Note on mutexes: The live-resource tracker mutex is confined to resource create/destroy and diagnostics. Retained track() calls use the resource's own shared-ownership state and do not acquire it. Fence and submission pools still synchronize their shared queue-level state.


4. Validation Cost Model

4.1 How It Works

Each validation check is tagged with one or more ValidationTag values (30+ tags covering CreateOp, DestroyOp, RenderCommandEncoderOp, CommandBufferOp, TextureOp, BufferOp, etc.). Tags are aggregated at build time into a constexpr uint64_t EnabledValidationTags bitmask.

The SNAP_RHI_VALIDATE macro uses if constexpr to test the bitmask. When a tag is disabled, the compiler removes the entire validation block — including the string literals, format arguments, and lambda captures used by the check.

4.2 Cost Summary

Build Configuration Validation Cost Binary Size Impact
Release (all tags OFF) Zero — code is not compiled in None
Debug, tags OFF Zero — same elimination None
Debug, selective tags ON Per-enabled-tag checks only Small
SNAP_RHI_ENABLE_ALL_VALIDATION All checks active + Vulkan layers + slow safety checks Significant — debug only

4.3 Fine-Grained Control

You don't have to choose between "all validation" and "none". Enable only the tags you need:

cmake -B build \
    -DSNAP_RHI_ENABLE_VULKAN=ON \
    -DSNAP_RHI_VALIDATION_RENDER_COMMAND_ENCODER_OP=ON \
    -DSNAP_RHI_VALIDATION_TEXTURE_OP=ON \
    -DCMAKE_BUILD_TYPE=Debug

This enables validation for render encoder and texture operations only — everything else compiles to zero cost.


5. Build Configurations for Performance Work

Goal Preset / Flags Notes
Profile (real-world perf) macos-metal-release / --release No validation, compiler optimizations enabled
Debug (correctness) macos-metal-demo Debug labels + logs, no validation overhead
Full validation macos-metal-demo-validation All checks enabled — expect slower execution
Selective validation Raw CMake with individual SNAP_RHI_VALIDATION_* flags Surgical debugging with minimal perf impact

Always profile in Release builds. Debug builds include debug labels, logging, and potentially unoptimized code that does not represent production performance. See the Profiling Guide for tool-specific setup.


6. Measuring Performance

6.1 GPU Timestamp Queries

SnapRHI provides cross-backend GPU timestamp queries via QueryPool:

auto queryPool = device->createQueryPool({ .queryCount = 2 });

commandBuffer->resetQueryPool(queryPool.get(), 0, 2);

// Timestamps are recorded at the encoding-scope boundaries.
auto* encoder = commandBuffer->getComputeCommandEncoder();
encoder->beginEncoding(snap::rhi::PassTimestampWrites{
    .queryPool = queryPool.get(), .beginningOfPassWriteIndex = 0, .endOfPassWriteIndex = 1});
// ... GPU work ...
encoder->endEncoding();

std::array<std::chrono::nanoseconds, 2> results{};
queryPool->getResults(0, 2, results);
auto gpuTime = results[1] - results[0];  // std::chrono::nanoseconds

6.2 Custom Profiling Labels

Enable SNAP_RHI_ENABLE_CUSTOM_PROFILING_LABELS to inject scoped markers:

SNAP_RHI_DECLARE_CUSTOM_PROFILING_LABEL(device, "Shadow Pass");
// ... rendering code — label is automatically popped at scope exit

These labels appear in platform GPU profilers (Xcode Instruments, RenderDoc, Nsight, AGI).

6.3 Platform Profiling Tools

Platform Recommended Tool What It Measures
macOS / iOS Xcode Instruments (Metal System Trace) GPU timeline, shader execution, memory bandwidth
Windows PIX, NVIDIA Nsight, RenderDoc GPU timeline, draw call breakdown, resource usage
Linux RenderDoc, NVIDIA Nsight Frame capture, pipeline state, GPU counters
Android AGI (Android GPU Inspector), Perfetto GPU counters, frame pacing, memory

Full details: Profiling Guide


Further Reading