Buffers
Buffer is a GPU memory allocation for storing typed data — uniforms, vertex data, index data, compute storage, or anything a shader needs to read or write.
Creating buffers (recommended)
For application-owned GPU memory, use Runtime and bind the returned Parcel in a scheme (with_parcel, set_vertex_buffer, MemoryExchange deposits). All Rust, Python, FFI, and .NET examples use this path.
#![allow(unused)] fn main() { use goldy::{BufferFlags, BufferKind}; let vertices = [/* Vertex2D ... */]; let vertex_parcel = runtime.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?; // Uninitialized storage (e.g. a uniform updated each frame via MemoryExchange deposit): let uniform = runtime.acquire_buffer_sized::<MyUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?; }
See runtime-owned-memory.md for textures, mosaics, and release.
With Raw Bytes
When the data is naturally &[u8], pass an explicit element stride to acquire_buffer:
#![allow(unused)] fn main() { use goldy::{BufferFlags, BufferKind}; // Stride defaults to 1 when omitted (byte-addressable) let parcel = runtime.acquire_buffer( raw_bytes.len() as u64, BufferKind::Scattered, None, BufferFlags::empty(), Some(&raw_bytes), )?; // Explicit stride for structured buffer views let parcel = runtime.acquire_buffer( raw_bytes.len() as u64, BufferKind::Scattered, Some(16), BufferFlags::empty(), Some(&raw_bytes), )?; // With flags (e.g. CPU_READABLE) let parcel = runtime.acquire_buffer( raw_bytes.len() as u64, BufferKind::Scattered, Some(16), BufferFlags::CPU_READABLE, Some(&raw_bytes), )?; }
Empty Buffer
#![allow(unused)] fn main() { let parcel = runtime.acquire_buffer( 4096, BufferKind::Scattered, None, BufferFlags::empty(), None, )?; // With a specific element stride let parcel = runtime.acquire_buffer( 4096, BufferKind::Scattered, Some(64), BufferFlags::empty(), None, )?; }
Low-level Runtime::alloc_* (crate-internal)
The runtime routes standalone allocations through VramAllocator via
crate-internal Runtime::alloc_buffer helpers. Application code should not call these;
use Runtime acquire APIs above.
Data Access Patterns
The access pattern describes how shader threads access the buffer. This drives hardware optimizations and determines the bindless descriptor category.
#![allow(unused)] fn main() { pub enum BufferKind { Scattered, // default — any thread, any address, read/write Broadcast, // all threads read the same address } }
| Pattern | Shader Mapping | Use When |
|---|---|---|
Scattered | StructuredBuffer<T>, RWStructuredBuffer<T> | General storage: particles, meshes, compute I/O |
Broadcast | ConstantBuffer / uniform buffer | Uniform data: transforms, time, settings |
For read-only input buffers that don't need write access, create with BufferKind::Scattered and access through goldy_buf_ro<T> in the shader. This enables hardware read-cache optimizations without requiring a separate access pattern.
BufferFlags
#![allow(unused)] fn main() { bitflags! { pub struct BufferFlags: u32 { const COPY_SRC = 1 << 0; const COPY_DST = 1 << 1; const CPU_READABLE = 1 << 2; const CPU_WRITABLE = 1 << 4; } } }
| Flag | Purpose |
|---|---|
COPY_SRC | Buffer can be a copy source |
COPY_DST | Buffer can be a copy destination |
CPU_READABLE | Placement hint: expect host claims on this parcel. Backends may keep the medium host-coherent so take() is wait + pointer. Semantics are identical without the flag (staged copy). |
CPU_WRITABLE | Host-mapped staging for deposits / upload copies. Prefer MemoryExchange::bind_deposit for application uploads. |
Query RuntimeCapabilities::has_zero_copy_storage_readback to detect whether the backend honors CPU_READABLE with a mapped pointer.
Writing Data
Prefer MemoryExchange::bind_deposit for CPU→GPU uploads. Direct host writes on CPU_WRITABLE staging parcels remain for deposit/staging internals:
Raw bytes
#![allow(unused)] fn main() { buffer.write(offset, &bytes)?; }
Typed data
#![allow(unused)] fn main() { buffer.write_data(offset, &[1.0f32, 2.0, 3.0])?; }
Both methods write at a byte offset from the start of the buffer.
Reading Data
Use a host claim after submit:
#![allow(unused)] fn main() { let mut submission = scheme.submit()?; let bytes = (&mut submission >> buffer.whole()).take::<u8>()?.to_vec(); }
Clearing
Zero-fill a region of the buffer:
#![allow(unused)] fn main() { buffer.clear(&device, offset, size)?; }
Bindless Descriptors
Every buffer with Scattered or Broadcast access is registered in the global bindless descriptor set. Schemes bind parcels via with_parcel; the opaque ResourceHandle is available for identity / retention checks:
#![allow(unused)] fn main() { // Opaque typed identity — equality / hashing only; no public heap index let handle = buffer.handle(ResourceAccess::Read).unwrap(); // Read-only SRV vs write UAV are distinct handles when both exist let srv_handle = buffer.handle(ResourceAccess::Read).unwrap(); }
BufferView
A BufferView is a sub-region of an existing Buffer with its own bindless descriptor. The shader sees the sub-region as a zero-based buffer.
Creating Views
#![allow(unused)] fn main() { // Raw byte view — offset, size, optional element stride let view = buffer.create_view(1024, 512, Some(16))?; // Typed view — first element index, element count let view = buffer.create_typed_view::<[f32; 4]>(0, 256)?; }
Using Views
Views implement BufferSource, so they work anywhere a Buffer does — set_vertex_buffer, set_index_buffer, write_data, clear, and scheme parcel binding:
#![allow(unused)] fn main() { let view_handle = view.handle(ResourceAccess::Read).unwrap(); pass.set_vertex_buffer(0, &view); }
Lifetime
Dropping a BufferView unregisters its descriptor but does not free the parent buffer's memory. Multiple views of the same buffer can exist simultaneously.
StructuredBufferElement
The StructuredBufferElement trait marks types safe for Runtime::acquire_buffer_with_data.
It is implemented for common multi-byte primitives (u16, u32, f32, f64, etc.), fixed-size arrays of those types, and #[repr(C)] structs via #[derive(goldy_derive::StructuredBufferElement)].
Not implemented for u8/i8 — passing &[u8] would set stride to 1, which almost never matches the shader's expected struct stride. Use Runtime::acquire_buffer with an explicit element stride for raw bytes.
Rust-generated Slang structs
#[derive(goldy::GpuType)] makes Rust the source of truth for a structured-buffer
element. Pass Type::GPU_TYPE when creating the shader and reference the type
without redeclaring it in authored Slang. Declare logical fields only — Goldy
packs to the Slang structured-buffer ABI at upload (acquire_buffer_with_data,
write_data) and injects reserved __goldy_padN fields in generated Slang.
#![allow(unused)] fn main() { #[repr(C)] #[derive(Copy, Clone, bytemuck::Pod, bytemuck::Zeroable, goldy::GpuType)] struct Particle { position: [f32; 3], color: [f32; 4], } let shader = ShaderModule::from_slang_with_gpu_types( &device, source, &[Particle::GPU_TYPE], )?; let particles = device.acquire_buffer_with_data(&particles, BufferKind::Scattered)?; }
// Particle is injected by Goldy.
[goldy_compute]
void cs_main(BufRO<Particle> particles, ThreadId id) {
Particle particle = particles[id.x];
}
Do not bytemuck::bytes_of a GpuType into a GPU buffer: that is the host
layout, not the storage ABI. Typed Goldy upload is the syscall.
The portable initial field set is f32, u32, i32, 2–4 lane arrays of those
scalars, and square f32 matrices. Unsupported or sub-word host fields fail with
an actionable error instead of silently changing the ABI.
Shared library modules can inject the same declarations via
ShaderLibrary::from_source_with_gpu_types. Prefer that over redeclaring the
struct in authored Slang. Stage I/O structs with semantics (POSITION,
TEXCOORD*, SV_Position) stay authored in the shader.
Matrix Convention
Goldy uses column-major matrix layout in uniform/constant buffers across all backends. Rust math libraries (glam, nalgebra, ultraviolet) already store matrices column-major, so upload directly without transposing:
#![allow(unused)] fn main() { let uniforms = MyUniforms { projection: proj.to_cols_array_2d(), modelview: view.to_cols_array_2d(), }; buffer.write_data(0, &[uniforms])?; }
Goldy sets SLANG_MATRIX_LAYOUT_COLUMN_MAJOR at the Slang session level, so DX12, Vulkan, and Metal all interpret float4x4 the same way.