Buffers

Buffer is a GPU memory allocation for storing typed data — uniforms, vertex data, index data, compute storage, or anything a shader needs to read or write.

For application-owned GPU memory, use Runtime and bind the returned Parcel in a scheme (with_parcel, set_vertex_buffer, MemoryExchange deposits). All Rust, Python, FFI, and .NET examples use this path.

#![allow(unused)]
fn main() {
use goldy::{BufferFlags, BufferKind};

let vertices = [/* Vertex2D ... */];
let vertex_parcel = runtime.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?;

// Uninitialized storage (e.g. a uniform updated each frame via MemoryExchange deposit):
let uniform = runtime.acquire_buffer_sized::<MyUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;
}

See runtime-owned-memory.md for textures, mosaics, and release.

With Raw Bytes

When the data is naturally &[u8], pass an explicit element stride to acquire_buffer:

#![allow(unused)]
fn main() {
use goldy::{BufferFlags, BufferKind};

// Stride defaults to 1 when omitted (byte-addressable)
let parcel = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    None,
    BufferFlags::empty(),
    Some(&raw_bytes),
)?;

// Explicit stride for structured buffer views
let parcel = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    Some(16),
    BufferFlags::empty(),
    Some(&raw_bytes),
)?;

// With flags (e.g. CPU_READABLE)
let parcel = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    Some(16),
    BufferFlags::CPU_READABLE,
    Some(&raw_bytes),
)?;
}

Empty Buffer

#![allow(unused)]
fn main() {
let parcel = runtime.acquire_buffer(
    4096,
    BufferKind::Scattered,
    None,
    BufferFlags::empty(),
    None,
)?;

// With a specific element stride
let parcel = runtime.acquire_buffer(
    4096,
    BufferKind::Scattered,
    Some(64),
    BufferFlags::empty(),
    None,
)?;
}

Low-level Runtime::alloc_* (crate-internal)

The runtime routes standalone allocations through VramAllocator via crate-internal Runtime::alloc_buffer helpers. Application code should not call these; use Runtime acquire APIs above.

Data Access Patterns

The access pattern describes how shader threads access the buffer. This drives hardware optimizations and determines the bindless descriptor category.

#![allow(unused)]
fn main() {
pub enum BufferKind {
    Scattered, // default — any thread, any address, read/write
    Broadcast, // all threads read the same address
}
}
PatternShader MappingUse When
ScatteredStructuredBuffer<T>, RWStructuredBuffer<T>General storage: particles, meshes, compute I/O
BroadcastConstantBuffer / uniform bufferUniform data: transforms, time, settings

For read-only input buffers that don't need write access, create with BufferKind::Scattered and access through goldy_buf_ro<T> in the shader. This enables hardware read-cache optimizations without requiring a separate access pattern.

BufferFlags

#![allow(unused)]
fn main() {
bitflags! {
    pub struct BufferFlags: u32 {
        const COPY_SRC      = 1 << 0;
        const COPY_DST      = 1 << 1;
        const CPU_READABLE  = 1 << 2;
        const CPU_WRITABLE  = 1 << 4;
    }
}
}
FlagPurpose
COPY_SRCBuffer can be a copy source
COPY_DSTBuffer can be a copy destination
CPU_READABLEPlacement hint: expect host claims on this parcel. Backends may keep the medium host-coherent so take() is wait + pointer. Semantics are identical without the flag (staged copy).
CPU_WRITABLEHost-mapped staging for deposits / upload copies. Prefer MemoryExchange::bind_deposit for application uploads.

Query RuntimeCapabilities::has_zero_copy_storage_readback to detect whether the backend honors CPU_READABLE with a mapped pointer.

Writing Data

Prefer MemoryExchange::bind_deposit for CPU→GPU uploads. Direct host writes on CPU_WRITABLE staging parcels remain for deposit/staging internals:

Raw bytes

#![allow(unused)]
fn main() {
buffer.write(offset, &bytes)?;
}

Typed data

#![allow(unused)]
fn main() {
buffer.write_data(offset, &[1.0f32, 2.0, 3.0])?;
}

Both methods write at a byte offset from the start of the buffer.

Reading Data

Use a host claim after submit:

#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;
let bytes = (&mut submission >> buffer.whole()).take::<u8>()?.to_vec();
}

Clearing

Zero-fill a region of the buffer:

#![allow(unused)]
fn main() {
buffer.clear(&device, offset, size)?;
}

Bindless Descriptors

Every buffer with Scattered or Broadcast access is registered in the global bindless descriptor set. Schemes bind parcels via with_parcel; the opaque ResourceHandle is available for identity / retention checks:

#![allow(unused)]
fn main() {
// Opaque typed identity — equality / hashing only; no public heap index
let handle = buffer.handle(ResourceAccess::Read).unwrap();

// Read-only SRV vs write UAV are distinct handles when both exist
let srv_handle = buffer.handle(ResourceAccess::Read).unwrap();
}

BufferView

A BufferView is a sub-region of an existing Buffer with its own bindless descriptor. The shader sees the sub-region as a zero-based buffer.

Creating Views

#![allow(unused)]
fn main() {
// Raw byte view — offset, size, optional element stride
let view = buffer.create_view(1024, 512, Some(16))?;

// Typed view — first element index, element count
let view = buffer.create_typed_view::<[f32; 4]>(0, 256)?;
}

Using Views

Views implement BufferSource, so they work anywhere a Buffer does — set_vertex_buffer, set_index_buffer, write_data, clear, and scheme parcel binding:

#![allow(unused)]
fn main() {
let view_handle = view.handle(ResourceAccess::Read).unwrap();
pass.set_vertex_buffer(0, &view);
}

Lifetime

Dropping a BufferView unregisters its descriptor but does not free the parent buffer's memory. Multiple views of the same buffer can exist simultaneously.

StructuredBufferElement

The StructuredBufferElement trait marks types safe for Runtime::acquire_buffer_with_data. It is implemented for common multi-byte primitives (u16, u32, f32, f64, etc.), fixed-size arrays of those types, and #[repr(C)] structs via #[derive(goldy_derive::StructuredBufferElement)].

Not implemented for u8/i8 — passing &[u8] would set stride to 1, which almost never matches the shader's expected struct stride. Use Runtime::acquire_buffer with an explicit element stride for raw bytes.

Rust-generated Slang structs

#[derive(goldy::GpuType)] makes Rust the source of truth for a structured-buffer element. Pass Type::GPU_TYPE when creating the shader and reference the type without redeclaring it in authored Slang. Declare logical fields only — Goldy packs to the Slang structured-buffer ABI at upload (acquire_buffer_with_data, write_data) and injects reserved __goldy_padN fields in generated Slang.

#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Copy, Clone, bytemuck::Pod, bytemuck::Zeroable, goldy::GpuType)]
struct Particle {
    position: [f32; 3],
    color: [f32; 4],
}

let shader = ShaderModule::from_slang_with_gpu_types(
    &device,
    source,
    &[Particle::GPU_TYPE],
)?;
let particles = device.acquire_buffer_with_data(&particles, BufferKind::Scattered)?;
}
// Particle is injected by Goldy.
[goldy_compute]
void cs_main(BufRO<Particle> particles, ThreadId id) {
    Particle particle = particles[id.x];
}

Do not bytemuck::bytes_of a GpuType into a GPU buffer: that is the host layout, not the storage ABI. Typed Goldy upload is the syscall.

The portable initial field set is f32, u32, i32, 2–4 lane arrays of those scalars, and square f32 matrices. Unsupported or sub-word host fields fail with an actionable error instead of silently changing the ABI.

Shared library modules can inject the same declarations via ShaderLibrary::from_source_with_gpu_types. Prefer that over redeclaring the struct in authored Slang. Stage I/O structs with semantics (POSITION, TEXCOORD*, SV_Position) stay authored in the shader.

Matrix Convention

Goldy uses column-major matrix layout in uniform/constant buffers across all backends. Rust math libraries (glam, nalgebra, ultraviolet) already store matrices column-major, so upload directly without transposing:

#![allow(unused)]
fn main() {
let uniforms = MyUniforms {
    projection: proj.to_cols_array_2d(),
    modelview: view.to_cols_array_2d(),
};
buffer.write_data(0, &[uniforms])?;
}

Goldy sets SLANG_MATRIX_LAYOUT_COLUMN_MAJOR at the Slang session level, so DX12, Vulkan, and Metal all interpret float4x4 the same way.