SnapRHI is designed for low-latency, high-throughput rendering workloads. Every abstraction is evaluated against the cost of the underlying API call it wraps. This document explains the architectural decisions that minimize overhead and how to verify performance in your own workloads.
Related: Profiling Guide | Resource Management | API Overview
- Design Principles
- Per-Backend Optimizations
- Overhead Budget
- Validation Cost Model
- Build Configurations for Performance Work
- Measuring Performance
Fixed-limit command operations use stack or inline storage. Variable-size state is reserved or pooled before normal recording and retains capacity across command-buffer reuse. It grows only when a workload exceeds that retained capacity; this is deliberately documented rather than hidden behind an absolute allocation-free claim.
| Backend | Hot-Path Allocation Strategy |
|---|---|
| Metal | Pooled ResourceBatcher instances use 256-entry inline buckets. Descriptor sets sub-allocate from a shared MTLBuffer. Large barrier lists are emitted in fixed-size chunks. |
| Vulkan | Fixed-limit descriptor, vertex-buffer, attachment, and clear-value translation uses stack/inline arrays. Large barrier lists are emitted in fixed-size chunks. Public command pools recycle native command storage. |
| OpenGL | A retained linear command arena placement-constructs commands contiguously. Batched copies and barriers are split into fixed-size command records. |
The OpenGL command allocator enforces this at the type level — all command
structs must be trivially destructible (verified by static_assert),
so the entire recording buffer can be rewound without walking individual destructors. The arena can grow if
CommandPoolCreateInfo::commandMemorySize is exhausted, then reuses the larger allocation on subsequent recordings.
Retained command buffers reserve common-case resource and interop-reference capacity at construction. An unusually
large recording can grow those vectors; CommandPoolCreateFlags::UnretainedResources removes shared-ownership and
storage work in Release builds when the application can guarantee lifetime safety itself.
Every validation check in SnapRHI is guarded by an if constexpr test against
a compile-time bitmask of enabled validation tags:
#define SNAP_RHI_VALIDATE(layer, condition, reportLevel, tags, ...) \
if constexpr ((static_cast<uint64_t>(tags) & snap::rhi::EnabledValidationTags) && \
(static_cast<uint64_t>(reportLevel) >= SNAP_RHI_ENABLED_REPORT_LEVEL)) { ...When a tag is disabled (the default for Release builds), the compiler
dead-code-eliminates the entire check — not branched over at runtime, but
absent from the binary. Over 30 individual validation tags can be toggled
independently, or all at once via SNAP_RHI_ENABLE_ALL_VALIDATION.
Release build: record → encode → submit (validation code is not in the binary)
Debug + validation: record → validate → encode → submit
See §4 — Validation Cost Model for details.
SnapRHI pools and recycles expensive objects rather than creating and destroying them per-frame:
| Resource | Pooling Mechanism |
|---|---|
| Fences | FencePool — acquires/releases reusable internal fences across all backends |
| Command buffers | Public CommandPool — preconstructs a bounded set of backend wrappers and recycles them across all backends |
| Command pools (Vulkan) | Each public pool owns one native VkCommandPool; retained and individually resettable command buffers reuse native allocations |
| Framebuffers (OpenGL) | FramebufferPool — caches GL framebuffer objects by hashed description |
| Descriptor memory (Metal) | Sub-allocates from a single MTLBuffer with atomic block tracking |
| Resource batchers (Metal) | Pooled ResourceBatcher instances with acquire/release lifecycle |
Applications can opt out of automatic reference retention for maximum
throughput. In Release builds, CommandPoolCreateFlags::UnretainedResources skips ResourceResidencySet
bookkeeping entirely. The application then guarantees every recorded resource's lifetime through GPU completion.
See Resource Management — Retention Modes for the full contract.
| Technique | Details |
|---|---|
| Descriptor sub-allocation | DescriptorPool sub-allocates from a single MTLBuffer with per-block atomic tracking — no per-set MTLBuffer creation. |
| Resource usage batching | ResourceBatcher aggregates useResource: calls with pre-sized buckets (256 initial), plus a _lastBucket cache for O(1) consecutive-add fast path. |
| Pooled batchers | Metal Device maintains a pool of ResourceBatcher instances — zero allocation during encoder setup. |
| Technique | Details |
|---|---|
| Application-sized command pools | Each public CommandPool preconstructs its bounded wrapper set and owns one native VkCommandPool. Native command buffers are reused except when an unretained, non-resettable buffer returns its allocation on release. |
| Descriptor pool allocation | Standard VkDescriptorPool block allocation — set creation is a fast pool sub-allocation. |
| Technique | Details |
|---|---|
| Linear command allocator | Arena-style allocator that grows by doubling (starting at 2 KB). Commands are POD structs, placement-new'd sequentially. Iterated via CommandIterator without any virtual dispatch. |
| Comprehensive state cache | StateCache tracks all GL state (buffers, textures, samplers, programs, framebuffers, viewport, depth, stencil, blending, vertex attributes, feature toggles). Every set*() call checks the cache first via inline should*() methods — redundant gl* calls are eliminated before reaching the driver. |
| Framebuffer cache | FramebufferPool caches GLuint framebuffer objects keyed by a hashed attachment description — identical render pass configurations reuse existing FBOs. |
| Command buffer pool | The public CommandPool uses the common fixed-capacity recycler; each reused buffer rewinds its CPU command arena instead of recreating it. |
No glGet* calls |
The state cache explicitly avoids glGet* queries during rendering, which would force a GPU→CPU sync and stall the pipeline. All state is tracked CPU-side. |
This section characterizes where SnapRHI adds overhead versus a raw backend call. The analysis is split by operation category and distinguishes between backends where behavior differs materially.
These operations are called many times per frame (inner-loop hot path).
| Operation | Metal | Vulkan | OpenGL |
|---|---|---|---|
draw / drawIndexed |
Virtual call → flushes deferred dynamic-offset/resource state when dirty → native draw | Virtual call → native draw | Virtual call → placement-new DrawCmd into arena. Actual GL calls happen later during replay. |
bindRenderPipeline |
Virtual call → static_cast → stores pipeline pointer + sets depth bias via Obj-C message → resourceResidencySet.track() |
Virtual call → static_cast → vkCmdBindPipeline → resourceResidencySet.track() |
Virtual call → placement-new SetRenderPipelineCmd into arena → resourceResidencySet.track() |
bindDescriptorSet |
Virtual call → binds argument-buffer state → stores bounded dynamic offsets → retains referenced resources | Virtual call → fixed stack array → vkCmdBindDescriptorSets |
Virtual call → placement-new BindDescriptorSetCmd + bounded dynamic-offset copy → retains referenced resources |
bindVertexBuffer |
Virtual call → native buffer bind → track() |
Virtual call → one-element stack translation → vkCmdBindVertexBuffers → track() |
Virtual call → placement-new SetVertexBufferCmd → track() |
bindVertexBuffers (multi) |
Virtual call → loop of native binds + track() |
Fixed stack arrays → vkCmdBindVertexBuffers → track() per buffer |
Virtual call → one bounded SetVertexBuffersCmd → track() per buffer |
bindIndexBuffer |
Virtual call → static_cast → stores pointer + offset (no native call) → track() |
Virtual call → static_cast → vkCmdBindIndexBuffer → track() |
Virtual call → placement-new SetIndexBufferCmd → track() |
setViewport |
Virtual call → struct conversion → [encoder setViewport:] |
Virtual call → static_cast → vkCmdSetViewport + vkCmdSetScissor |
Virtual call → placement-new SetViewportCmd. During replay: gl.viewport() through state cache shouldViewport() |
setDepthBias |
Virtual call → stores values → conditional [encoder setDepthBias:...] |
Virtual call → static_cast → vkCmdSetDepthBias |
Virtual call → placement-new SetDepthBiasCmd |
setStencilReference |
Virtual call → stores values → [encoder setStencilFrontReferenceValue:...] |
Virtual call → static_cast → vkCmdSetStencilReference |
Virtual call → placement-new SetStencilReferenceCmd |
setBlendConstants |
Virtual call → [encoder setBlendColorRed:...] |
Virtual call → static_cast → vkCmdSetBlendConstants |
Virtual call → placement-new SetBlendConstantsCmd |
- Metal and Vulkan call native APIs inline during encoding. Metal retains some dirty resource/dynamic-offset state for a bounded flush before draw or dispatch; Vulkan binds descriptor sets immediately.
- OpenGL records all commands into a linear allocator (arena). The actual
gl*calls happen duringreplayCommands()at submit time, where theRenderCmdPerformerwalks the arena viaCommandIterator(no virtual dispatch during replay — just aswitchon command IDs). - Fixed API limits such as
MaxVertexBuffers,MaxBoundDescriptorSets, andMaxAttachmentssize translation arrays, avoiding per-call heap allocation while making invalid ranges explicit precondition violations.
Every bind/record operation calls resourceResidencySet.track() per resource.
In retained mode (default), this obtains a strong reference directly from the resource's weak_from_this() state
and appends it to the command buffer's pre-reserved vector. It does not touch the diagnostic live-resource tracker and
therefore performs no tracker mutex acquisition or hash lookup. Consecutive duplicate resources are rejected before
the weak lock.
In unretained mode (CommandPoolCreateFlags::UnretainedResources), track() returns after its mode check in
Release builds. In debug builds with slow validations, recording stores weak_ptr observations and submission
snapshots them for fence-scoped lifetime checking instead.
| Retention Mode | Per-track() Cost |
|---|---|
| Retained (default) | Weak-to-strong reference acquisition + pre-reserved vector::push_back; adjacent duplicates are a pointer comparison |
| Unretained (Release) | One mode check; no reference-count or container work |
| Unretained (Debug + slow validations) | Weak lifetime observation appended to pre-reserved diagnostic storage |
Implication: Retained mode is lock-free but still pays shared ownership bookkeeping. Use
CommandPoolCreateFlags::UnretainedResourcesonly when the application can enforce lifetimes through GPU completion.
| Phase | Metal | Vulkan | OpenGL |
|---|---|---|---|
beginEncoding (RenderPass) |
Creates MTLRenderPassDescriptor → native encoder → tracks attachments |
Tracks attachments → resolves inline image views/clear values → vkCmdBeginRenderPass |
Placement-new BeginRenderPassCmd into arena → tracks attachments |
beginEncoding (Dynamic Rendering) |
Creates MTLRenderPassDescriptor → native encoder |
Tracks attachments → builds an inline bounded attachment list → vkCmdBeginRendering |
Placement-new BeginRenderPass1Cmd into arena |
endEncoding |
[encoder endEncoding] → clears pipeline state |
vkCmdEndRenderPass or vkCmdEndRendering → resets descriptor state |
Placement-new EndRenderPassCmd |
Vulkan attachment and clear-value translations are bounded by public API limits and remain inline. Any native image layout work remains part of the backend operation rather than a container-allocation cost.
| Step | Cost | Notes |
|---|---|---|
submissionTracker.tryReclaim() |
Mutex lock + linear scan of in-flight submissions | Reclaims completed slots; bounded by max in-flight count |
| Texture interop processing | Per-interop-texture work | Only when interop textures are present |
Fence acquisition (via buildFence) |
FencePool::acquireFence() — mutex lock + linear scan for reusable fence |
Pool avoids createFence on hot path |
| Semaphore wait/signal | Per-semaphore native wait/signal | Metal: Obj-C dispatch. Vulkan: native semaphore handles. OpenGL: glFenceSync |
| Native submit | [mtlCommandBuffer commit] / vkQueueSubmit / GL command replay |
Core driver cost |
submissionTracker.track() |
Retained: mutex lock + strong command-buffer/fence resolution + slot update. Slow validation: weak lifetime snapshot + slot update | Retained mode, or unretained mode only when slow validation is enabled |
OpenGL-specific: replayCommands() |
Reset GL state cache → iterate arena via CommandIterator (switch dispatch) → RenderCmdPerformer calls gl* through state cache |
Queue use is externally synchronized by contract; no internal queue mutex is taken. State-cache comparisons filter redundant native calls. |
All creation methods go through Device::createResource<>(), which:
new— heap-allocates and constructs the backend implementation object, including its native API initialization.- Wraps it in
std::shared_ptrwith a custom deleter that carries its live-resource tracker slot. - Registers the object in the live-resource tracker — mutex lock + amortized vector append at a new high-water mark, or O(1) reuse of a vacant slot. Reused slots require no allocation.
The tracker stores one raw pointer and an embedded free-list link per slot. Resource deleters unregister before object
destruction, allowing diagnostic snapshots to obtain weak_from_this() under the tracker mutex without storing a
second weak reference per resource.
| Operation | Additional Backend Cost |
|---|---|
| Create texture | Metal: [device newTextureWithDescriptor:]. Vulkan: vkCreateImage + vkAllocateMemory + vkBindImageMemory + image view creation. OpenGL: glGenTextures + glTexStorage*. |
| Create buffer | Metal: [device newBufferWithLength:]. Vulkan: vkCreateBuffer + vkAllocateMemory + vkBindBufferMemory. OpenGL: glGenBuffers + glBufferData. |
| Create render pipeline | Metal: [device newRenderPipelineStateWithDescriptor:] (may compile shaders). Vulkan: vkCreateGraphicsPipelines (may compile shaders). OpenGL: glLinkProgram + state translation. |
| Create sampler | Metal: [device newSamplerStateWithDescriptor:]. Vulkan: vkCreateSampler. OpenGL: glGenSamplers + glSamplerParameter*. |
| Create descriptor set | Metal: argument buffer sub-allocation from DescriptorPool. Vulkan: vkAllocateDescriptorSets from pool. OpenGL: stores binding metadata. |
| Create command buffer | Recycles a backend wrapper from a CommandPool; normal shared_ptr/live-tracker bookkeeping still applies. Metal: fresh one-shot [commandQueue commandBuffer]. Vulkan: retained or individually resettable pools reset and reuse the same native command buffer; unretained pools without individual reset allocate after the previous native buffer was freed on release. OpenGL: reuses its reserved CommandAllocator arena. |
Pipeline creation is the most expensive operation — it may trigger shader compilation. Use
PipelineCacheto amortize this cost.
Destruction happens when the last shared_ptr reference is released:
- The custom deleter unregisters its known slot — mutex lock + O(1) free-list insertion, with no lookup or allocation.
- (Slow validations only) Scans incomplete submissions for weak observations of the erased resource.
- The backend object is destroyed, or a command-buffer wrapper is released to its pool. In an unretained Vulkan pool without individual reset, release also returns the native command-buffer allocation.
Descriptor writes (bindUniformBuffer, bindTexture, bindSampler,
updateDescriptorSet) are virtual calls into backend-specific logic:
- Metal: Writes into argument buffer memory (sub-allocated from
DescriptorPool). No native API call — direct memory writes. - Vulkan: Calls
vkUpdateDescriptorSetswith translated write structs. - OpenGL: Stores binding metadata into internal arrays (applied during replay).
- Metal:
[buffer contents]+ offset — returns pointer to shared memory. - Vulkan:
vkMapMemory/vkUnmapMemory— may involve driver synchronization. - OpenGL:
glMapBufferRange/glUnmapBuffer.
No additional SnapRHI overhead beyond the virtual call + static_cast.
| Build Configuration | Cost |
|---|---|
| Tag-gated validation disabled (Release default) | Zero for those checks — if constexpr removes them; unrecoverable critical invariant reporting remains compiled |
| Validation enabled (selective tags) | Per-enabled-tag checks at each guarded call site |
Slow validations (SNAP_RHI_ENABLE_SLOW_VALIDATIONS) |
Additional per-track() weak observations + submission snapshot + lifetime scans when resources are destroyed |
| Operation Category | Overhead vs. Raw API | Dominant Cost |
|---|---|---|
| Draw / dispatch | Virtual call plus native recording/replay work | Metal flushes dirty bounded state; Vulkan records directly; OpenGL writes its arena |
| Dynamic state (viewport, depth bias, etc.) | Virtual call | Metal/Vulkan: inline native call. OpenGL: arena write (replay adds state cache check) |
| Bind vertex/index buffer | Virtual call + native/recorded bind + track() |
Retained mode: lock-free shared ownership; fixed-limit translations remain inline |
| Bind descriptor set | Virtual call + backend bind/record | Exact dynamic-offset validation is compiled out when disabled; fixed arrays cover API limits |
| Begin/end render pass | Virtual call + attachment setup | Fixed-limit translations remain inline; native descriptor/layout work still applies |
| Submit command buffer | Fence pool + submission tracker | Submission/fence trackers synchronize shared state; OpenGL performs full command replay |
| Resource creation | shared_ptr + live-tracker registration |
Mutex lock + amortized slot-vector growth + native API object creation. Pipelines: may compile shaders |
| Resource destruction | shared_ptr custom deleter |
Mutex lock + O(1) slot recycle + native API release |
| Descriptor writes | Virtual call + backend write | Metal: memory write. Vulkan: vkUpdateDescriptorSets. OpenGL: store metadata |
Note on virtual dispatch: SnapRHI uses virtual functions for backend polymorphism. On modern CPUs, a predicted virtual call costs ~1–2 ns. This is the cost of a uniform, type-safe multi-backend API. Applications that build against a single backend can expect the compiler to devirtualize
finalclass methods in many cases (all backend classes are markedfinal).Note on mutexes: The live-resource tracker mutex is confined to resource create/destroy and diagnostics. Retained
track()calls use the resource's own shared-ownership state and do not acquire it. Fence and submission pools still synchronize their shared queue-level state.
Each validation check is tagged with one or more ValidationTag values (30+
tags covering CreateOp, DestroyOp, RenderCommandEncoderOp,
CommandBufferOp, TextureOp, BufferOp, etc.). Tags are aggregated at
build time into a constexpr uint64_t EnabledValidationTags bitmask.
The SNAP_RHI_VALIDATE macro uses if constexpr to test the bitmask.
When a tag is disabled, the compiler removes the entire validation block —
including the string literals, format arguments, and lambda captures used
by the check.
| Build Configuration | Validation Cost | Binary Size Impact |
|---|---|---|
| Release (all tags OFF) | Zero — code is not compiled in | None |
| Debug, tags OFF | Zero — same elimination | None |
| Debug, selective tags ON | Per-enabled-tag checks only | Small |
SNAP_RHI_ENABLE_ALL_VALIDATION |
All checks active + Vulkan layers + slow safety checks | Significant — debug only |
You don't have to choose between "all validation" and "none". Enable only the tags you need:
cmake -B build \
-DSNAP_RHI_ENABLE_VULKAN=ON \
-DSNAP_RHI_VALIDATION_RENDER_COMMAND_ENCODER_OP=ON \
-DSNAP_RHI_VALIDATION_TEXTURE_OP=ON \
-DCMAKE_BUILD_TYPE=DebugThis enables validation for render encoder and texture operations only — everything else compiles to zero cost.
| Goal | Preset / Flags | Notes |
|---|---|---|
| Profile (real-world perf) | macos-metal-release / --release |
No validation, compiler optimizations enabled |
| Debug (correctness) | macos-metal-demo |
Debug labels + logs, no validation overhead |
| Full validation | macos-metal-demo-validation |
All checks enabled — expect slower execution |
| Selective validation | Raw CMake with individual SNAP_RHI_VALIDATION_* flags |
Surgical debugging with minimal perf impact |
Always profile in Release builds. Debug builds include debug labels, logging, and potentially unoptimized code that does not represent production performance. See the Profiling Guide for tool-specific setup.
SnapRHI provides cross-backend GPU timestamp queries via QueryPool:
auto queryPool = device->createQueryPool({ .queryCount = 2 });
commandBuffer->resetQueryPool(queryPool.get(), 0, 2);
// Timestamps are recorded at the encoding-scope boundaries.
auto* encoder = commandBuffer->getComputeCommandEncoder();
encoder->beginEncoding(snap::rhi::PassTimestampWrites{
.queryPool = queryPool.get(), .beginningOfPassWriteIndex = 0, .endOfPassWriteIndex = 1});
// ... GPU work ...
encoder->endEncoding();
std::array<std::chrono::nanoseconds, 2> results{};
queryPool->getResults(0, 2, results);
auto gpuTime = results[1] - results[0]; // std::chrono::nanosecondsEnable SNAP_RHI_ENABLE_CUSTOM_PROFILING_LABELS to inject scoped markers:
SNAP_RHI_DECLARE_CUSTOM_PROFILING_LABEL(device, "Shadow Pass");
// ... rendering code — label is automatically popped at scope exitThese labels appear in platform GPU profilers (Xcode Instruments, RenderDoc, Nsight, AGI).
| Platform | Recommended Tool | What It Measures |
|---|---|---|
| macOS / iOS | Xcode Instruments (Metal System Trace) | GPU timeline, shader execution, memory bandwidth |
| Windows | PIX, NVIDIA Nsight, RenderDoc | GPU timeline, draw call breakdown, resource usage |
| Linux | RenderDoc, NVIDIA Nsight | Frame capture, pipeline state, GPU counters |
| Android | AGI (Android GPU Inspector), Perfetto | GPU counters, frame pacing, memory |
Full details: Profiling Guide
- Profiling Guide — Platform-specific profiling tools and workflows
- Resource Management — Lifetime rules, retention modes, memory patterns
- API Overview — Full API reference and design philosophy
- Debugging Guide — Validation layers, sanitizers, diagnostics