Goldy Logo

Goldy: GPU runtime for the Fondaco Machine

Goldy is a Rust GPU library that realizes the Fondaco Machine.

Maturity: Goldy 0.3 is the Fondaco Machine public API. SemVer applies within 0.3.x; expect breaking changes at 0.4.

The Fondaco model in Goldy

FondacoGoldy
SchemeScheme — retained graph, resubmitted each frame
ParcelParcel, Buffer, Texture
DispatchCompute, render, copy, and present nodes inside a scheme
ExchangeSurfaceExchange, MemoryExchange
SettlementTransaction → Claim

Goldy abstracts where bytes live (descriptor slots, residency, relocation) but exposes what access costs (access patterns, resize cost, readback path). See the design thesis.

Typed bindless shaders

Shaders use Slang with goldy_exp virtual entry points. Resources are typed parameters — Goldy resolves bindless slots automatically:

import goldy_exp;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(MyUniforms cfg, Scattered<uint> data, ThreadId id) {
    data[id.x] = data[id.x] + cfg.base;
}
TypeUse
Scattered<T>Read/write storage
BufRO<T>Read-only storage
DirectSpatial<T>Read/write texture
Interpolated<T>Sampled texture
Broadcast (struct param)Per-dispatch constants

Scheme

Record once, submit every frame. Goldy inserts barriers, parallelizes independent nodes, and aliases transient resources:

#![allow(unused)]
fn main() {
let mut scheme = Scheme::new(&ctx);
scheme
    .node("simulate", &sim_pipeline)
    .with_parcel(&particles, NodeAccess::ReadWrite)
    .dispatch(group_count, 1, 1);
let submission = scheme.submit()?;
}

Compute-to-surface

Compute shaders can write swapchain drawables directly — no graphics pipeline or raster pass:

#![allow(unused)]
fn main() {
let surface = SurfaceExchange::new(&ctx, &window, SurfaceConfig::default())?;
let mut scheme = Scheme::new(&ctx);
let (lease, present) = surface.bind_destination(&mut scheme)?;
scheme
    .node("render", &compute_pipeline)
    .with_parcel(&uniforms, NodeAccess::Read)
    .with_present(&lease)
    .dispatch(wg_x, wg_y, 1);

let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

Backends and bindings

PlatformBackend
WindowsDX12 (default), Vulkan
LinuxVulkan (Wayland surfaces)
macOSMetal

CUDA and WebGPU backends are in progress. A Tenstorrent backend (Torus) is planned. See Backend Architecture.

Bindings: Python, .NET, C++, Rust FFI Client.

License

MIT — see License.

Installation

Requirements

Adding Goldy to Your Project

[dependencies]
goldy = "0.2"

Or with cargo:

cargo add goldy

Feature Flags

FeatureDefaultDescription
vulkanyesVulkan 1.4+ backend (Linux, Windows); implies graphics and gpu
dx12yesDirectX 12 backend (Windows); implies graphics and gpu
metalyesMetal Tier 2+ backend (macOS); implies graphics and gpu
graphicsyesRaster pipelines, render targets, surfaces, and presentation
tensoryesDense tensor algebra over parcels (independent of graphics; not an ML framework)
gpuyesImplied by every real GPU backend (not mock). Do not enable alone.
cudanoCUDA backend (in progress; NVIDIA compute; implies gpu, not graphics)
webgpunoWebGPU backend (in progress; via wgpu; implies graphics and gpu)
instrumentationyesStructured tracing via tracing-subscriber (zero-cost when disabled)

graphics is implied by the native backends. gpu is implied by every real GPU backend (not mock). Textures and samplers remain available without graphics for GPGPU workloads. CUDA is not a platform default; it auto-selects only in --no-default-features --features cuda builds (otherwise set GOLDY_BACKEND=cuda):

cargo test --no-default-features --features cuda --test scheme_compute_integration

Platform-inappropriate features are no-ops — enabling metal on Linux or dx12 on macOS compiles cleanly but does nothing.

To build with only specific backends:

[dependencies]
goldy = { version = "0.2", default-features = false, features = ["vulkan"] }

Shader Toolchain

Goldy uses Slang as its shader language. The Rust build.rs downloads (if needed) and embeds the pinned Slang version at compile time; at runtime Goldy extracts and loads it automatically. Application developers do not install Slang separately.

Set GOLDY_SLANG_PATH only to override with a custom Slang build.

Verifying Installation

use goldy::{RuntimeDescriptor, Instance, RequestAdapterOptions};

fn main() -> anyhow::Result<()> {
    let instance = Instance::new()?;

    println!("Available GPUs:");
    for adapter in instance.enumerate_adapters() {
        println!("  {} ({:?})", adapter.name, adapter.device_type);
    }

    let device = instance
        .request_adapter(&RequestAdapterOptions::default())?
        .request_runtime(&RuntimeDescriptor::default())?;
    println!("\nUsing: {}", device.adapter_info().name);

    Ok(())
}
cargo run

Expected output:

Available GPUs:
  NVIDIA GeForce RTX 4060 Ti (DiscreteGpu)
  Intel(R) UHD Graphics 770 (IntegratedGpu)

Using: NVIDIA GeForce RTX 4060 Ti

Backend Selection

Goldy selects the best backend for your platform automatically:

PlatformDefault Backend
WindowsDX12
LinuxVulkan
macOSMetal

Override at runtime with GOLDY_BACKEND:

GOLDY_BACKEND=vulkan cargo run

Platform-Specific Setup

Windows

DX12 is used by default and requires no additional setup. For the Vulkan backend, install the Vulkan SDK. Ensure your GPU drivers are up to date.

Linux

Install Vulkan development packages:

# Ubuntu/Debian
sudo apt install libvulkan-dev vulkan-tools

# Fedora
sudo dnf install vulkan-loader-devel vulkan-tools

# Arch
sudo pacman -S vulkan-icd-loader vulkan-tools

macOS

Goldy uses the native Metal backend — no MoltenVK or Vulkan SDK needed. Ensure macOS 12+ and Xcode command-line tools are installed:

xcode-select --install

Windowing (for examples)

Examples require the examples feature (winit is gated):

cargo run --features examples --example triangle --release

Next Steps

Your First Triangle

This tutorial draws a colored triangle in a window using Goldy's render pipeline and present-on-scheme API (SurfaceExchange + Scheme + Transaction).

See examples/triangle.rs for the full source.

Recording the Scheme

Once at init (and again on resize), record a retained scheme: offscreen render pass → copy to surface via SurfaceExchange::bind_render_target.

#![allow(unused)]
fn main() {
use goldy::{
    shader::builtins, Buffer, BufferKind, Color, RuntimeDescriptor, Instance, Lease, LeaseRenderTarget,
    NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, Transaction, Vertex2D,
};

fn record_scheme(
    scheme: &mut Scheme,
    surface: &SurfaceExchange,
    pipeline: &RenderPipeline,
    vertex_buffer: &Buffer,
    scene_rt: &Lease<LeaseRenderTarget>,
    bg_color: Color,
) -> anyhow::Result<Transaction> {
    let mut pass = scheme.render_pass("triangle", scene_rt, TargetLoad::Clear(bg_color));
    pass.with_parcel(vertex_buffer, NodeAccess::Read);
    pass.set_pipeline(pipeline);
    pass.set_vertex_buffer(0, vertex_buffer);
    pass.draw(0..3, 0..1);
    pass.finish();
    surface.bind_render_target(scheme, scene_rt).map_err(Into::into)
}
}

Per-Frame Submit

Each frame submits the retained scheme and consumes the surface claim:

#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

Walkthrough

Instance, Runtime, and Context

#![allow(unused)]
fn main() {
let instance = Instance::new()?;
let device = Arc::new(
    instance
        .request_adapter(&RequestAdapterOptions::default())?
        .request_runtime(&RuntimeDescriptor::default())?,
);
let ctx = device.create_context()?;
}

Instance discovers available GPUs. create_context opens the submission context used by Scheme.

Vertex Buffer

#![allow(unused)]
fn main() {
let vertices = [
    Vertex2D::new(0.0, -0.5, Color::RED),
    Vertex2D::new(-0.5, 0.5, Color::GREEN),
    Vertex2D::new(0.5, 0.5, Color::BLUE),
];
let vertex_buffer = device.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?;
}

Vertex2D is a built-in vertex type with position and color. The runtime owns the retained parcel for the buffer's lifetime.

Shader and Pipeline

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang(&device, builtins::VERTEX_COLOR_2D)?;
let surface = SurfaceExchange::new(&ctx, window.as_ref(), SurfaceConfig::default())?;
let pipeline = RenderPipeline::new(
    &device,
    &shader,
    &shader,
    &RenderPipelineDesc {
        vertex_layout: Vertex2D::layout(),
        target_format: surface.format(),
        ..Default::default()
    },
)?;
}

builtins::VERTEX_COLOR_2D uses [goldy_vertex] and [goldy_fragment] virtual entry points from the goldy_exp library.

Surface and Presentation

#![allow(unused)]
fn main() {
let mut scheme = Scheme::new(&ctx);
let scene_rt = ctx.lease_render_target(width, height, surface.format(), None)?;
let present = record_scheme(&mut scheme, &surface, &pipeline, &vertex_buffer, &scene_rt, bg_color)?;

// Each frame:
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

SurfaceExchange manages the OS swapchain. Scene color is rendered to a scheme-leased offscreen target, copied to the drawable, and displayed when the claim is consumed. Rendering stays on the GPU — no CPU readback.

On resize, rebuild the scheme and transaction with the new dimensions (see examples/triangle.rs).

Run It

cargo run --example triangle

You should see a window with a colored triangle on a dark blue background.

Next Steps

Your First Compute Shader

This tutorial renders an animated plasma effect by dispatching a compute shader directly to a swapchain drawable — no graphics pipeline, no vertex buffers, no render passes.

See examples/compute_to_surface.rs for the full source.

The Shader

The compute shader uses goldy_exp virtual entry points. It reads uniforms via BufRO<Uniforms> and writes pixels via DirectSpatial<float4>:

import goldy_exp;

struct Uniforms {
    uint width;
    uint height;
    float time;
};

[goldy_compute]
[numthreads(8, 8, 1)]
void cs_main(BufRO<Uniforms> uniforms_buf, DirectSpatial<float4> output, ThreadId tid) {
    Uniforms u = uniforms_buf[0];

    if (tid.x >= u.width || tid.y >= u.height)
        return;

    float2 uv = float2(float(tid.x) / float(u.width),
                       float(tid.y) / float(u.height));
    float2 p = uv * 2.0 - 1.0;
    p.x *= float(u.width) / float(u.height);

    float t = u.time;
    float v = 0.0;
    v += sin(p.x * 6.0 + t);
    v += sin(p.y * 6.0 + t * 1.3);
    v += sin((p.x + p.y) * 4.0 + t * 0.7);
    v += sin(length(p) * 8.0 - t * 2.0);
    v *= 0.25;

    float3 col = float3(0.5 + 0.5 * sin(v * 3.14159 + 0.0),
                        0.5 + 0.5 * sin(v * 3.14159 + 2.094),
                        0.5 + 0.5 * sin(v * 3.14159 + 4.188));
    output[tid.xy] = float4(col, 1.0);
}

Key points:

  • BufRO<Uniforms> is a read-only structured buffer. Index with [0] to load the single element.
  • DirectSpatial<float4> is an RWTexture2D<float4> — write to it with output[tid.xy].
  • ThreadId maps to SV_DispatchThreadID. Each thread handles one pixel.
  • The [goldy_compute] attribute tells the Goldy compiler to wire up bindless slots automatically.

Rust Side

Uniform Buffer

Define the uniform struct on the Rust side with matching layout:

#![allow(unused)]
fn main() {
#[goldy::gpu]
struct Uniforms {
    width: u32,
    height: u32,
    time: f32,
}
}

Create the buffer with BufferKind::Scattered so it gets a bindless descriptor:

#![allow(unused)]
fn main() {
let uniform_buffer = device.acquire_buffer_with_data(
    &[Uniforms { width, height, time: 0.0 }],
    BufferKind::Scattered,
)?;
}

Compute Pipeline and Scheme

Compile the Slang source, create a ComputePipeline, and record a retained scheme once via SurfaceExchange::bind_destination:

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang(&device, COMPUTE_SHADER)?;
let compute_pipeline = ComputePipeline::new(&device, &shader)?;

let surface = SurfaceExchange::new_with_depth(&ctx, window.as_ref(), 3, SurfaceConfig::default())?;

let mut scheme = Scheme::new(&ctx);
let wg_x = width.div_ceil(8);
let wg_y = height.div_ceil(8);
let (lease, present) = surface.bind_destination(&mut scheme)?;
scheme
    .node("compute", &compute_pipeline)
    .with_parcel(&uniform_buffer, NodeAccess::Read)
    .with_present(&lease)
    .dispatch(wg_x, wg_y, 1);
}

Rendering a Frame

Each frame: upload new uniform values via a small upload scheme with a bound deposit ((&deposit << &uniforms)?), submit the main scheme, then present with (&mut submission >> &present).take()?.

#![allow(unused)]
fn main() {
fn render_frame(state: &mut RenderState) -> Result<()> {
    let (width, height) = state.surface.size();
    let elapsed = state.start_time.elapsed().as_secs_f32();

    let uniforms = Uniforms {
        width,
        height,
        time: elapsed,
    };

    (&state.uniform_deposit << &uniforms)?;
    state.upload_scheme.submit()?;

    let mut submission = state.scheme.submit()?;
    (&mut submission >> &state.present).take()?;
    Ok(())
}
}

At init, bind the deposit once on a retained upload scheme:

#![allow(unused)]
fn main() {
let mut upload_scheme = Scheme::new(&ctx);
let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
    &mut upload_scheme,
    goldy::DepositTarget::buffer_elements::<Uniforms>(&uniform_buffer, 1),
)?;
}

Step by Step

Update uniforms — MemoryExchange::bind_deposit records the upload topology once; each frame tender (&deposit << &uniforms)? before the main submit. write remains for offsets and partial fills.

Record the scheme once — SurfaceExchange::bind_destination registers the present exchange and returns a PresentLease plus a Transaction. scheme.node() creates a compute node bound to a pipeline. with_parcel() declares the uniform buffer dependency. with_present() binds the drawable lease. dispatch() sets the workgroup count.

Submit and present — scheme.submit() records and submits GPU work. (&mut submission >> &present).take()? presents the swapchain image. The explicit &mut borrow is required by operator semantics and leaves other claims on the submission untouched. The compute shader already wrote the pixels — there is no blit or copy step.

Run It

cargo run --example compute_to_surface

You should see an animated plasma pattern filling the window, rendered entirely from compute.

Next Steps

Parcels

A parcel is the unit of data Goldy schemes actually operate on: a whole buffer, a range within a buffer, or a texture. Every resource you acquire from a Runtime or a transient allocator hands you one or more parcels, and every with_parcel call on a scheme node passes exactly one.

#![allow(unused)]
fn main() {
use goldy::{BufferKind, ResourceAccess};

let parcel = runtime.acquire_buffer_with_data(&particles, BufferKind::Scattered)?;
let handle = parcel.handle(ResourceAccess::Write).unwrap();
let again = parcel.handle(ResourceAccess::Write).unwrap();
assert_eq!(handle, again);
}

You never declare layouts, allocate descriptor pools, or manage binding slots yourself. You acquire a parcel, bind it to a node with an access mode, and Goldy figures out the rest at dispatch time.

Categories

Every parcel belongs to one of five categories, matching the shape of access a shader can perform on it:

CategoryWhat it holdsShader-side type
ScatteredRead/write structured dataScattered<T>
BufRORead-only structured dataBufRO<T>
BroadcastSmall uniform data shared by every invocationa plain struct parameter
InterpolatedSampled texture dataInterpolated<T>
DirectSpatialRead/write texture dataDirectSpatial<T>
FilterSampler stateFilter

A parcel's category is fixed when it's created (BufferKind::Scattered, BufferKind::Broadcast, etc.) and determines which shader-side type it can satisfy. Categories are also independent identity spaces: a Scattered parcel and a Broadcast parcel are unrelated even if they happen to occupy "slot 3" internally — that internal indexing is not something client code ever sees or reasons about.

ResourceHandle and ResourceAccess

ResourceAccess (Read, Write, ReadWrite) describes the kind of access a piece of shader-visible data supports — for example, whether a buffer is exposed to the shader as read-only or read/write. parcel.handle(access) returns a ResourceHandle: an opaque, comparable identity for that parcel/access pair.

ResourceHandle is intentionally opaque. You can compare two handles for equality (useful for deciding whether a retained scheme needs to be re-recorded after a resource was reallocated), but there's nothing else to extract from one — it's an identity, not a number you're meant to interpret.

This is distinct from NodeAccess (Read, Write, ReadWrite, Overwrite), which is what you actually pass to with_parcel. NodeAccess describes how a scheme node uses a parcel for scheduling and hazard tracking (including Overwrite for "I'm replacing this data wholesale, don't preserve prior contents"); ResourceAccess is the narrower, resolved access a shader parameter requires.

Binding Parcels to Schemes

You bind parcels to compute or render nodes with with_parcel, in the same order the shader declares its resource parameters:

#![allow(unused)]
fn main() {
scheme
    .node("update", &pipeline)
    .with_parcel(&params_buf, NodeAccess::Read)
    .with_parcel(&particle_buf, NodeAccess::ReadWrite)
    .dispatch((particle_count + 63) / 64, 1, 1);
}

At dispatch time, Goldy checks each bound parcel's category against what the shader's reflected signature expects. If slot 0 expects Broadcast (from the shader's SimParams params parameter) but you bound a Scattered parcel there, binding fails with a clear error instead of silently producing garbage or undefined behavior.

Typed Resource Parameters in Shaders

On the shader side, goldy_exp provides types that mirror the categories above and map directly to underlying Slang resource types. These appear as ordinary parameters on virtual entry points:

Goldy TypeUnderlying Slang TypeUsage
Scattered<T>RWStructuredBuffer<T>Read/write buffer: data[i], data[i].field = v
BufRO<T>StructuredBuffer<T>Read-only buffer: buf[i]
Interpolated<T>Texture2D<T>Sampled texture: tex.Sample(samp, uv)
DirectSpatial<T>RWTexture2D<T>Writable texture: img[int2(x,y)]
ByteAddressRWByteAddressBufferRaw byte access: .Load(), .Store(), .Interlocked*()
FilterSamplerStateSampler for texture filtering

Any user-defined struct type (e.g. MyUniforms) declared as a parameter is automatically treated as Broadcast — no wrapper type needed.

Contrast with Traditional Binding

Traditional (Vulkan/DX12)Goldy Parcels
SetupDeclare descriptor set layouts, allocate pools, create and update descriptor setsAcquire a parcel; category is fixed at creation
BindingBind descriptor sets before each draw/dispatchPass parcels via with_parcel on scheme nodes
Shader accesslayout(set=0, binding=1) buffer ...Scattered<T> data as a function parameter
ValidationRuntime errors or silent corruption on mismatchCategory checks at dispatch time
Cross-backendLayout declarations differ per APISame shader code on Vulkan, DX12, and Metal

Example: Compute Shader with Parcels

Shader (particle_update.slang):

import goldy_exp;

struct SimParams {
    float dt;
    uint count;
};

struct Particle {
    float2 pos;
    float2 vel;
};

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(SimParams params, Scattered<Particle> particles, ThreadId id) {
    if (id.x >= params.count) return;

    Particle p = particles[id.x];
    p.pos += p.vel * params.dt;
    particles[id.x] = p;
}

Rust dispatch:

#![allow(unused)]
fn main() {
let params_buf = runtime.acquire_buffer_with_data(&[sim_params], BufferKind::Broadcast)?;
let particle_buf = runtime.acquire_buffer_with_data(&particles, BufferKind::Scattered)?;

let shader = ShaderModule::from_slang(&device, PARTICLE_UPDATE_SOURCE)?;
let pipeline = ComputePipeline::new(&device, &shader)?;

let mut scheme = Scheme::new(&ctx);
scheme
    .node("update", &pipeline)
    .with_parcel(&params_buf, NodeAccess::Read)
    .with_parcel(&particle_buf, NodeAccess::ReadWrite)
    .dispatch((particle_count + 63) / 64, 1, 1);
scheme.submit()?;
}

The shader author writes natural function parameters. The Rust side binds parcels in declaration order via with_parcel. Everything below that — slot packing, descriptor heaps, cross-backend plumbing — is an implementation detail you never need to think about.

Schemes

A scheme is Goldy's retained unit of work: a recorded graph of dispatches and precedences you submit again without re-recording while it stays clean. Create one with Scheme::new, bind parcels on nodes, and call submit every frame. Compute pipelines and record-time constant buffers are interned on the scheme, so you can drop the objects you used only to record.

#![allow(unused)]
fn main() {
let mut scheme = Scheme::new(&ctx);
scheme
    .node("double", &pipeline)
    .with_parcel(&data, NodeAccess::ReadWrite)
    .dispatch(1, 1, 1);
scheme.submit()?;
}

Structural mutation (new nodes, new bindings, include) drops retained command lists. Params-only mutation (pipeline, scalars, dispatch dims) re-records only the partitions whose baked payload changed. A clean scheme resubmits with neither recording nor fingerprint hashing.

Nesting

Scheme::include copies a child's description into the parent as one group. The copy is a snapshot: mutating the child afterward does not change the parent, and the child remains independently submittable.

#![allow(unused)]
fn main() {
let mut attn = Scheme::new(&ctx);
attn.node("rmsnorm", &rmsnorm).with_parcel(&x, NodeAccess::ReadWrite).dispatch(groups, 1, 1);
// ... more child nodes ...

let mut layer = Scheme::new(&ctx);
layer.include("attn", &attn)?.finish();
layer.submit()?;
attn.submit()?; // still legal
}

Scheme::group is sugar for a temporary child on the same context plus include:

#![allow(unused)]
fn main() {
layer.group("ffn", |s| {
    s.node("up", &up).with_parcel(&x, NodeAccess::ReadWrite).dispatch(groups, 1, 1);
    Ok(())
})?;
}

GroupBuilder::after adds a group-level precedence (A completes before B). Record order is the total order: only forward precedences are admitted. Expansion to node pairs happens when the schedule cache is rebuilt, not on the clean submit path.

Included nodes keep group provenance. Backend markers and validation text show the path (layer0/attn/rmsnorm).

Include restrictions

Anything not proven correct is rejected at include time (GoldyError::Validation, or StaleResource if a child stamp is dead). The parent is left untouched. v1 admits only:

  • The same Context (cross-context include is rejected)
  • No pending record errors on the child
  • No cpu_node (closures and per-occurrence staging are instance state)
  • No deposits, present leases, or swapchain outputs (exchanges belong to the submitting root)
  • No yielding dispatches
  • No transient parcels (lease-epoch semantics across two submitters are not yet admitted)

Interleaving parent and child submits on shared parcels is correct: the ledger orders them like any two schemes. The parent will topology_dirty and re-record barriers — correct, not free.

Virtual Entry Points

Goldy's virtual entry points let you write shader entry points with clean, typed parameters instead of raw uniform uint slots and SV_* semantics. You annotate your function with [goldy_compute], [goldy_vertex], or [goldy_fragment], and a source-to-source transform generates the real Slang [shader("...")] entry point with all the bindless plumbing wired up.

The Attributes

AttributeStageGenerated Slang Attribute
[goldy_compute]Compute[shader("compute")]
[goldy_vertex]Vertex[shader("vertex")]
[goldy_fragment]Fragment[shader("fragment")]
[goldy_raygen]Ray generation[shader("raygeneration")]
[goldy_miss]Miss[shader("miss")]
[goldy_closesthit]Closest hit[shader("closesthit")]
[goldy_mesh]Mesh[shader("mesh")]
[goldy_amplification]Amplification / task[shader("amplification")]

A minimal example:

import goldy_exp;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<uint> data, ThreadId id) {
    data[id.x] = data[id.x] * 2;
}

This is equivalent to manually writing a [shader("compute")] entry point with uniform uint push-constant parameters, descriptor heap lookups, and SV_DispatchThreadID — but without any of that boilerplate.

What Virtual Entry Points Accept

Resource Parameters

Each resource parameter occupies one bindless slot (a 16-bit index packed into push constants). The generated wrapper calls the corresponding goldy_* free function to resolve the slot to a live GPU handle.

Parameter TypeResolves ViaDescription
Scattered<T>goldy_scattered<T>(slot)Read/write storage buffer
BufRO<T>goldy_buf_ro<T>(slot)Read-only storage buffer
Interpolated<T>goldy_interpolated<T>(slot)Sampled 2D texture
DirectSpatial<T>goldy_direct_spatial<T>(slot)Read/write 2D texture
ByteAddressgoldy_byte_address(slot)Raw byte-address buffer
Filtergoldy_filter(slot)Sampler state
Accelgoldy_accel(slot)Top-level acceleration structure (RayQuery / TraceRay, when RuntimeCapabilities::ray_query or ray_tracing_pipelines is set)

Broadcast Parameters

Any user-defined struct type that isn't a recognized resource or system-value type is treated as a broadcast (constant buffer). The generated code calls goldy_broadcast<T>(slot) to fetch the entire struct from a uniform buffer:

struct SimParams { float dt; uint count; };

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(SimParams params, Scattered<Particle> data, ThreadId id) {
    // params is fetched from a constant buffer automatically
}

In vertex and fragment shaders, the last unrecognized struct is treated as the stage input (vertex attributes or fragment varyings) rather than a broadcast. All preceding unrecognized structs are broadcasts. Payload structs are shader-owned: define them once (or with matching semantics in each module) and Goldy links producer → consumer structurally. Rust never repeats a Varying type.

System-Value Parameters

System-value wrapper types are mapped to SV_* semantics. The generated entry point declares the raw semantic parameter and constructs the wrapper:

Wrapper TypeMaps ToAvailable Fields
ThreadIdSV_DispatchThreadID.x, .y, .z, .xy, .xyz
GroupThreadIdSV_GroupThreadID.x, .y, .z, .xy, .xyz
GroupIdSV_GroupID.x, .y, .z, .xy, .xyz
VertexIdSV_VertexID.value
InstanceIdSV_InstanceID.value
IsFrontFaceSV_IsFrontFace.value
DispatchRaysIndexSV_DispatchRaysIndex.x, .y, .z (raygen)
DispatchRaysDimensionsSV_DispatchRaysDimensions.x, .y, .z (raygen)

Scalar Parameters

Plain scalar types (uint, float, int, bool, and vector variants) become user parameters — full-precision u32 words in a separate region of the push constants. Bind them with with_param on the compute node builder (after with_parcel calls for resource params):

#![allow(unused)]
fn main() {
scheme
    .node("offset", &pipeline)
    .with_parcel(&data, NodeAccess::ReadWrite)
    .with_param(offset)
    .dispatch(1, 1, 1);
}

Scalar params are the only facts the runtime can bake. A value in a Scattered buffer or a broadcast struct is a parcel read; Goldy cannot see the word, so it cannot specialize on it. A mode flag, loop bound, or feature toggle passed with with_param is already on the CPU. If it holds still across clean submits, retained-scheme prediction recompiles the dispatch with that word as a preprocessor literal — a real Slang + driver compile, not a constant patch — and the driver can delete whatever that constant makes unreachable. Gate expensive work behind those scalars (if (has_tint != 0u) { ... }) rather than behind a load from a constant buffer if that elision is the point. Authors do not opt sites in; putting the fact in the right place is the whole contract. Details, including when this is worth expecting, are in Shader Specialization Prediction.

Pass-Through Parameters

In vertex and fragment shaders, the last unrecognized struct parameter passes through as a stage input (vertex attributes or interpolated varyings). It appears directly in the generated entry point signature without bindless resolution:

[goldy_fragment]
float4 fs_main(MyUniforms cfg, FullscreenVarying input) : SV_Target {
    // cfg → broadcast (slot 0)
    // input → pass-through stage input (interpolated varyings)
    return float4(cfg.time, 0, 0, 1);
}

The Source-to-Source Transform

The transform (implemented in slang/virtual_main.rs) runs before Slang compilation and performs three operations:

  1. Generates a wrapper function with the real [shader("...")] attribute and a fixed 16-word push-constant signature (mesh wrappers also declare out vertices / out indices).
  2. Renames the user function to _goldy_user_<name> (or goldy_mesh / goldy_amplification for those stages) so both can coexist.
  3. Removes the [goldy_*] attribute and [numthreads] / [outputtopology] from the renamed helper (they live on the generated wrapper).

Push Constant Layout

The generated entry point always declares a fixed signature regardless of how many parameters the user function has:

Words  0–7:  _bw0.._bw7   — 16 × u16 bindless indices packed 2 per word
Words  8–15: _uw0.._uw7   — 8 × u32 user scalar parameters

Bindless indices are packed as pairs into 32-bit words: the low 16 bits of _bw0 hold slot 0, the high 16 bits hold slot 1, and so on. This fits up to 16 resource/broadcast parameters and 8 scalar parameters in 64 bytes of push constants.

Before and After

What you write:

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(TimeUniforms cfg, Scattered<uint> data, ThreadId id) {
    data[id.x] = data[id.x] + cfg.base;
}

What gets compiled (generated wrapper prepended, user function renamed):

[shader("compute")]
[numthreads(64, 1, 1)]
void cs_main(uniform uint _bw0, ..., uniform uint _bw7,
             uniform uint _uw0, ..., uniform uint _uw7,
             uint3 _sv0 : SV_DispatchThreadID) {
    TimeUniforms cfg = goldy_broadcast<TimeUniforms>(_bw0 & 0xFFFFu);
    Scattered<uint> data = goldy_scattered<uint>((_bw0 >> 16u) & 0xFFFFu);
    ThreadId id = ThreadId(_sv0);
    _goldy_user_cs_main(cfg, data, id);
}

// Original function, renamed:
void _goldy_user_cs_main(TimeUniforms cfg, Scattered<uint> data, ThreadId id) {
    data[id.x] = data[id.x] + cfg.base;
}

The #line 1 directive is inserted between the generated wrapper and the user source so that compiler diagnostics report correct line numbers.

Vertex/Fragment Example

Define the interpolated payload once. Each stage lists only the resources it reads. Goldy links VSOutput to the fragment input by semantic (SV_Position, TEXCOORD0, …), not by struct name, and merges scene into one pipeline slot.

struct VSOutput {
    float4 position : SV_Position;
    float2 uv       : TEXCOORD0;
};

[goldy_vertex]
VSOutput vs_main(SceneUniforms scene, Scattered<Instance> instances, VertexId vid, InstanceId iid) {
    Instance inst = instances[iid.value];
    VSOutput out;
    // ... transform vertex ...
    return out;
}

[goldy_fragment]
float4 fs_main(SceneUniforms scene, Interpolated<float4> albedo, Filter samp,
               VSOutput input) : SV_Target {
    return albedo.Sample(samp, input.uv) * scene.tint;
}

On the CPU, bind by those parameter names:

#![allow(unused)]
fn main() {
pass.with_shader_bindings(&[
    ShaderBinding::read("scene", &scene),
    ShaderBinding::read("instances", &instances),
    ShaderBinding::read("albedo", &albedo),
    ShaderBinding::sampler("samp", &sampler),
]);
pass.set_pipeline(&pipeline);
}

The merged raster contract is fragment-first (scene, albedo, samp) then unique vertex resources (instances). Wrappers remap each stage onto that shared push layout, so the vertex shader still sees scene as its local first parameter while reading pipeline slot 0.

Builtins (VertexId, InstanceId) are invocation-provided; they are not part of the resource contract.

Preprocessor Conditionals

Virtual entry points support #ifdef/#else/#endif blocks directly inside the parameter list. This is useful for shader variants like MSAA:

[goldy_compute]
[numthreads(4, 16, 1)]
void cs_main(BufRO<uint> config,
#ifdef msaa
             BufRO<uint> mask_lut, DirectSpatial<float4> out_tex,
#else
             DirectSpatial<float4> out_tex,
#endif
             ThreadId tid) {
    // ...
}

The transform generates conditional blocks in the wrapper's signature, body, and call arguments so that the correct branch is selected at compile time based on preprocessor defines.

Rust Compute Kernels

Goldy can lower a restricted Rust GPU dialect into canonical [goldy_compute] Slang at compile time, then prepare and record through the normal Scheme path.

This is the initial design for issue #78. It is not arbitrary Rust, a second runtime compiler, or CUDA <<<>>> syntax. Slang remains the runtime backend compiler; the proc-macro is an AOT frontend that produces structured KernelDef metadata and typed record helpers.

To step the same kernel on the CPU without a handwritten Rust twin, see CPU host-callable shaders (issue #292).

Quick example

#![allow(unused)]
fn main() {
use goldy::gpu;

#[goldy::compute(workgroup_size = [256, 1, 1])]
fn saxpy(x: &[f32], y: &mut [f32], a: f32) {
    let i = gpu::global_id().x;
    if i < y.len() {
        y[i] = a * x[i] + y[i];
    }
}

// Host:
let kernel = saxpy::Kernel::prepare(&device)?;
kernel
    .record(&mut scheme, "saxpy", &x, &y, a)
    .over_1d(n);
// or exact grid counts:
kernel
    .record(&mut scheme, "saxpy", &x, &y, a)
    .groups([n.div_ceil(256), 1, 1]);
}

prepare compiles (or hits the shader cache) once. record only appends Scheme topology — it does not launch into a stream. Use use goldy::gpu; (or goldy::gpu::global_id()) for builtins.

Signature mapping

Rust GPU-dialect types use the same names as shaders/goldy_exp/access.slang (BufRO, Scattered, DirectSpatial, ThreadId, …).

Rust parameterSlang / Scheme
&[T] / gpu::BufRO<T>BufRO<T>, NodeAccess::Read
&mut [T]Scattered<T>, NodeAccess::ReadWrite
gpu::Scattered<T>Scattered<T>, NodeAccess::Write
gpu::Tensor<T>parent BufRO<T> + packed layout, NodeAccess::Read
gpu::TensorMut<T>parent Scattered<T> + packed layout, NodeAccess::ReadWrite
gpu::TensorWrite<T>parent Scattered<T> + packed layout, NodeAccess::Write
gpu::Uniform<T>broadcast resource, NodeAccess::Read
gpu::DirectSpatial<gpu::Float4>DirectSpatial<float4>, NodeAccess::Write (swapchain lease or texture)
u32 / i32 / f32 / booltyped scalar push words (no manual to_bits)

Hidden builtins (appended to the Slang signature when used):

RustSlang
gpu::global_id()ThreadId
gpu::local_id()GroupThreadId
gpu::workgroup_id()GroupId
gpu::workgroup_barrier()GroupMemoryBarrierWithGroupSync
let mut s = gpu::workgroup_array::<T, N>()file-scope groupshared T s[N]
gpu::workgroup_sum::<N>(val, scratch)tree-reduce sum; every lane gets the total
gpu::workgroup_max::<N>(val, scratch)tree-reduce max; every lane gets the max
gpu::workgroup_softmax_in_place::<N>(buf, base, count, scratch)in-place softmax over buf[base .. base+count)

workgroup_size is fixed on the attribute / KernelDef. .groups / .over_* only control the grid. A different workgroup size is a different pipeline.

Workgroup arrays are a fixed size known at compile time (not dynamic shared memory). Declare them at the kernel top level, then index them like a buffer.

workgroup_sum / workgroup_max / workgroup_softmax_in_place are 1D collectives. N must be a power of two (typically workgroup_size.x). They return the reduced value to every lane and include a trailing barrier, so the result is immediately usable. Softmax writes buf[base + t] for t < count; unused lanes contribute identity (-1e30 / 0). All threads in the workgroup must execute the call (no divergent branches around it). workgroup_sum/workgroup_max must be a let or simple assignment, not nested in a larger expression. Omit ::<N> to use workgroup_size.x. When buf is a tensor parameter, softmax indexes through the view (logical base + t).

Logical tensors vs physical buffers

gpu::Tensor<T> / TensorMut<T> / TensorWrite<T> bind a TensorView: the shader receives the parent parcel plus a scheme-owned packed metadata parcel (GoldyTensorLayout per tensor, one buffer for the dispatch). Indexing is logical-view-relative:

  • view[i] delinearizes i through rank/shape, then applies offset and strides
  • view.len() is the logical numel
  • view.dim(axis) and view.rank() read checked layout facts

Ordinary &[T] / &mut [T] / gpu::Scattered<T> stay the physical-index escape hatch: buf[i] is an element index in the parent buffer, and .len() is the buffer length. Tensor record methods take TensorView arguments, validate dtype, writeability, and optional shape contracts, and return Result because packing the layout (or a contract miss) can fail.

Shape contracts

Annotate tensor parameters with #[tensor(shape = [...])] to fix rank and to require equal extents at record time. Dimensions may be _ (any extent), an integer literal (exact), or an identifier (symbolic equality within one record call):

#![allow(unused)]
fn main() {
#[goldy::compute(workgroup_size = [256, 1, 1])]
fn rmsnorm(
    #[tensor(shape = [dim])] x: gpu::Tensor<f32>,
    #[tensor(shape = [dim])] weight: gpu::Tensor<f32>,
    #[tensor(shape = [dim])] out: gpu::TensorWrite<f32>,
) { /* ... */ }
}

Unannotated tensor parameters keep today's any-shape behavior. The generated record method checks every tensor argument against the contract before binding parcels or appending GraphIR. Failures name the kernel, parameter, axis, expected spec, and actual shape. Shader parameter order, the 48-byte GoldyTensorLayout ABI, and KERNEL_ABI_VERSION are unchanged.

Relationships that are not dimension equality — for example query-head / KV-head divisibility — stay explicit kernel or domain checks, not part of this DSL.

Goldy only has eight user scalar words, so layouts are not push constants. The metadata parcel is interned on the scheme, read-only in GraphIR, and does not need an external TensorKernels keepalive.

Architecture

Rust kernel
    │
    ▼
goldy_derive::compute
    ├── syn AST validation (GPU dialect)
    ├── goldy_shader_ir
    ├── canonical [goldy_compute] Slang
    └── KernelDef / KernelParam ABI
    │
    ▼
Kernel::prepare(device)
    └── existing ShaderModule + ComputePipeline + cache
    │
    ▼
typed record() → SchemeNodeBuilder bindings in declaration order

Raw hand-written [goldy_compute] shaders continue to work. Simple sources can also be parsed into the same KernelDef shape via goldy::slang::try_kernel_def_from_source, and wrappers can be emitted from ABI metadata with emit_wrapper_from_kernel_def so both paths share frame-table / PushLayout lowering.

Supported dialect (MVP)

Allowed: scalar arithmetic/comparisons, let / let mut, assignment, field/index access, if/else, while, for i in 0..n, casts, selected math intrinsics (abs/min/max/floor/ceil/sqrt/sin/cos/exp/pow/length), vector constructors (gpu::float2/float3/float4), buffer .len(), tensor .len() / .dim(axis) / .rank(), return, workgroup shared arrays + barriers, workgroup sum/max/softmax collectives, and the ID builtins above.

#[goldy::gpu] structs may be passed as &[T] uniforms; prepare prepends the generated Slang struct.

Rejected with span diagnostics: allocation, iterators/closures, traits/dyn, recursion, async, panics, arbitrary std calls, usize/isize, references except resource parameters, and unsupported patterns.

Element types for buffer slices are currently u32 / i32 / f32 / bool, or a #[goldy::gpu] struct for read-only &[T].

Diagnostics and dumps

  1. Proc-macro errors at Rust compile time for unsupported syntax.
  2. Slang / pipeline errors during Kernel::prepare.
  3. Backend errors after target compilation.

Set GOLDY_DUMP_RUST_KERNELS=1 (or a directory path) to dump canonical Slang and ABI metadata at prepare time.

Out of scope (later)

Graphics stages, GpuType derive, CUDA scalar parity polish, dynamic shared memory, specialization, and broad Rust compatibility belong to later phases / the wider goldy-jit roadmap.

CPU Dispatches

A CPU dispatch is a scheme node whose body runs on the host instead of the GPU. It is the "virtual main" idea from Virtual Entry Points applied to a plain Rust function: the parameter list is the entry point. Every buffer parcel you bind arrives as a whole &[T] or &mut [T] slice, followed by any scalar parameters.

#![allow(unused)]
fn main() {
use goldy::{NodeAccess, Scheme};

scheme
    .cpu_node("integrate")
    .with_parcel(&velocities, NodeAccess::Read)
    .with_parcel(&positions, NodeAccess::ReadWrite)
    .with_param(dt.to_bits())
    .dispatch(|vel: &[f32], pos: &mut [f32], dt: f32| {
        for (p, v) in pos.iter_mut().zip(vel) {
            *p += v * dt;
        }
    })?;
}

CPU dispatches exist so a host program with a Fondaco shape — a set of functions over parcels with declared access — can move into a scheme one node at a time. Each node can later be rewritten as a wave-based compute dispatch without touching its neighbours, because the scheme sees the same parcels and the same access declarations either way.

The virtual main

Any Fn + Send + Sync + 'static with up to sixteen CpuArg parameters is a valid main:

Parameter typeBound byNotes
&[T] where T: bytemuck::Podwith_parcel(.., NodeAccess::Read)whole parcel, byte_size / size_of::<T>() elements
&mut [T] where T: bytemuck::Podwith_parcel(.., Write / ReadWrite / Overwrite)same
u32, i32, f32, boolwith_param(u32)wire word; f32 via to_bits()

Slice parameters come first, in with_parcel order, then scalars in with_param order. dispatch validates the function against the bindings at record time and fails without recording anything when the arity, mutability, or element size does not match.

There is no thread id and no workgroup. The function runs once per submission and sees the complete parcel. It must not hold mutable state between submissions (it is Fn, not FnMut); everything it needs comes through its parameters.

Access and staging

Host visibility is a property of the node, not of the parcels. Bound parcels keep their device-resident allocation; the runtime stages them around the host call:

NodeAccessBefore the callSlice contentsAfter the call
Readdevice → host copycurrent parcel bytesnothing
Write, ReadWritedevice → host copycurrent parcel byteshost → device copy
Overwritenothingzeroedhost → device copy

Use Overwrite when the function produces every element; use Write when it touches only some of them and the rest must keep their previous values.

Host claims ((&mut submission >> &parcel).take()) follow the same medium rule: the parcel stays device-resident; mapped backends expose a coherent pointer after a timeline wait, others copy through a context staging pool. See Settlement.

Because the staging is a fence wait, a CPU dispatch is a full pipeline drain: every GPU node it depends on has finished before it runs, and every GPU node that depends on it starts only after its upload copy. A scheme with a CPU dispatch in the middle costs at least two extra GPU submits and one host wait per submission. This is the intended price of the migration path, not a steady-state design; on unified-memory backends a later pass may skip the copies without changing the node's contract.

What stays the same

  • Ordering. CPU dispatches take part in the same conflict analysis as GPU nodes. Two CPU dispatches on disjoint parcels are independent (they still run serially on the host); a CPU dispatch that reads a parcel written by a compute node runs after it.
  • Cross-scheme sync. Bound parcels are stamped like any other binding, so other schemes and contexts see the host's writes through the normal ledger.
  • Retention. A clean scheme with CPU dispatches still resubmits without re-recording; the GPU partitions around the host node are retained as usual. The host partition itself is never retained.
  • Leases. with_lease binds a context-minted buffer lease the same way with_parcel binds a retained parcel.

Textures are not supported as CPU dispatch parameters in 0.2.x.

Yielding Scripts

A yielding script is a compute shader whose lanes may suspend — hand a request to the host (or to another dispatch), let the runtime service it, and resume later with the answer. In Fondaco terms the lane petitions the runtime at a yield point; the runtime resolves the petition and re-enters the script in a continuation.

Nothing about a GPU wave can actually pause, so the runtime implements this the only way a GPU allows: the suspend point is a dispatch boundary. A lane that yields appends a small record (its petition payload and whatever state it wants back) to a mailbox and returns. After the dispatch retires, the runtime services the mailbox and launches the continuation over the recorded lanes. You write the two halves as ordinary Slang functions; the lowering, the mailboxes, the read-backs, and the resume dispatches are Goldy's job.

import goldy_exp;

[goldy_petition(Result = BufRO<uint>)]
struct Fetch { uint key; };

struct St { uint lane; uint acc; };

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<uint> data, uint scale, ThreadId tid) {
    uint v = data[tid.x];
    if (v % 2u == 1u) {
        $yield(cs_resume, Fetch { v }, St { tid.x, v * scale });
        return;
    }
    data[tid.x] = v * 2u;
}

[goldy_resume]
[numthreads(32, 1, 1)]
void cs_resume(Scattered<uint> data, Resolved<uint> r, St s, ThreadId tid) {
    data[s.lane] = r.is_null() ? 0xFFFFFFFFu : r[0] + s.acc;
}
#![allow(unused)]
fn main() {
use goldy::{NodeAccess, Petition, Promised, Scheme, YieldPoint};

#[repr(C)]
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
struct Fetch { key: u32 }

impl Petition for Fetch {
    const SLANG_NAME: &'static str = "Fetch";
    type Result = u32;
}

let node = scheme
    .node("fetch", &pipeline)
    .with_parcel(&data, NodeAccess::ReadWrite)
    .with_param(3)
    .yield_point(
        "cs_resume",
        YieldPoint::cpu(1024, 4096, |p: &Fetch, promised: Promised<'_, u32>| {
            match table.get(&p.key) {
                Some(value) => promised.fulfil(&[*value]),
                None => promised.reject(),
            }
        }),
    )
    .dispatch(16, 1, 1);
scheme.submit()?;
let stats = scheme.yield_stats(node).unwrap();
}

The scheme sees one node. Inside it the runtime runs as many rounds as the script needs.

dispatch(16, 1, 1) is the prologue launch: 16 groups × [numthreads(64,1,1)] = 1024 lanes, which matches the mailbox capacity. They are chosen independently — capacity is “how many lanes may be suspended at once”, not a dispatch size — and Backpressure::Stall chunks a wider prologue so the live population never exceeds capacity. arena_len (4096) is a third knob: how many result elements all fulfilments of one round may occupy, not a lane count.

The Slang side

Three constructs. import goldy_exp pulls in Resolved<T> and the petition helpers when Goldy sets GOLDY_YIELD — automatically, for yielding scripts and for GPU handlers that call goldy_resolve / goldy_reject:

ConstructMeaning
[goldy_petition(Result = BufRO<E>)] struct P { .. };P is a petition payload; a handler answers it with zero or more E elements.
$yield(continuation, payload, state);Suspend this lane. payload is a P, state is any struct the continuation wants back. Follow it with return;.
[goldy_resume] void c(program params.., Resolved<E> r, S s [, ThreadId tid])A continuation. Runs once per resumed lane.

Resolved<E> is a read-only window into the yield point's result arena: r.is_null() is true when the handler rejected the petition, r.len() is the element count, and r[i] reads an element. Indexing a null view is undefined.

The [goldy_compute] entry is the prologue. Continuations take a subset of its program parameters, matched by name and type — that is how the runtime knows which of your bindings to hand to the resume dispatch. payload and state may be written as P { a, b }, { a, b }, or any expression of the right type. A continuation may $yield again, to itself or to another continuation; the petition type of a continuation is inferred from the payloads yielded to it, or spelled out as [goldy_resume(P)].

Continuations declare their own [numthreads(N, 1, 1)] (default 64). Only ThreadId is available in a continuation and it counts resumed records, not the original lanes; carry anything you need from the prologue in the state struct.

Restrictions in v0

  • Petition and state structs hold only uint / int / float, fixed arrays of those, and nested structs of the same shape. This is what lets the host size mailboxes without a reflection round-trip; the Rust Petition type is then a #[repr(C)] mirror.
  • Program parameters are buffers (Scattered<T>, BufRO<T>, broadcast structs) and scalars. No textures or samplers yet, and no #if inside a parameter list.
  • One [goldy_compute] entry per source, and $yield must appear directly in the prologue or a continuation body, not in a helper.
  • Indirect dispatch (dispatch_shape_parcel) is not supported for yielding nodes.
  • goldy_buf_len is not portable: on Metal, WebGPU, and CUDA it currently returns 0xFFFFFFFF. Pass an explicit uint count (or a constant) for bounds checks and table wraps; do not use key % goldy_buf_len(table) in a handler.

The host side

ComputePipeline::new on a yielding script compiles the prologue and one entry point per continuation; ComputePipeline::is_yielding() reports it. Recording differs from a plain dispatch in one call: every continuation needs a yield_point(name, YieldPoint) before dispatch. Missing, duplicate, or mistyped yield points are reported on submit as GoldyError::Validation.

A YieldPoint has three parts:

  • capacity — the mailbox size: how many lanes may be suspended at this continuation at once.
  • arena_len — how many E elements all fulfilments of one round may use together.
  • a handler — YieldPoint::cpu(..) runs a Rust closure once per petition on the submitting thread; YieldPoint::node(.., &pipeline) runs a compute dispatch instead.

CPU handlers

#![allow(unused)]
fn main() {
YieldPoint::cpu(capacity, arena_len, |p: &P, promised: Promised<'_, E>| { .. })
}

The closure receives each payload as &P and a Promised<E>. fulfil(&[E]) copies the elements into the arena and the continuation sees them through Resolved<E>; reject() (or dropping the Promised) resumes the lane with a null view. A fulfilment that does not fit the arena is treated as a rejection and counted in YieldStats::arena_overflow.

The runtime checks P::SLANG_NAME against the continuation's petition and size_of::<P>() against the Slang struct at record time.

GPU handlers

#![allow(unused)]
fn main() {
YieldPoint::node(capacity, arena_len, &handler_pipeline).with_parcel(&table, NodeAccess::Read)
}

The handler shader is an ordinary [goldy_compute] entry with the signature

void cs_main(BufRO<P> petitions, Scattered<Resolution> resolutions, Scattered<E> arena,
             /* parcels from with_parcel, in order */ uint count, ThreadId tid)

dispatched over count lanes along x. It must write one resolution per petition with goldy_resolve(resolutions, i, offset, len) or goldy_reject(resolutions, i), placing results in arena itself. Nothing in the round touches the host.

Backpressure

What happens when more lanes yield than capacity:

PolicyBehaviour
Backpressure::Stall (default)Never lose a lane. The prologue is launched in chunks of at most capacity lanes, each chunk drained before the next starts. Since every lane yields at most once per body, the live population never exceeds the smallest Stall capacity. Needs capacity >= numthreads.x, and — when more than one chunk is needed — a ThreadId parameter on the prologue, no GroupId / GroupThreadId, and a one-dimensional dispatch.
Backpressure::DropLaunch once; lanes that find the mailbox full write nothing and their continuation never runs. Counted in YieldStats::dropped.

Statistics

Scheme::yield_stats(node) returns the counters of the last submission: chunks, rounds, petitions serviced, lanes resumed, rejections, drops, and arena overflows. They are what the tests assert on and a cheap way to see whether a capacity is sized right.

Execution model and cost

The yielding node is recorded as a host-driven node with the user's bindings for graph ordering. On each submission the driver:

  1. clears the yield counters and launches the prologue (or one chunk of it), writing mailbox set A;
  2. reads back the counters and the payloads of CPU-handled mailboxes;
  3. services every pending continuation — CPU handlers deposit their resolution table and arena bytes, GPU handlers are recorded as a dispatch;
  4. launches each pending continuation over its records, reading set A and writing any re-yields into set B;
  5. repeats from 2 with the sets swapped until no lane is suspended.

Each round is one sub-scheme submission on the same context, with a fence wait in the middle. Like a CPU dispatch, a yielding node is therefore a full pipeline drain, and the round count — not the lane count — is what costs. Scripts that yield once are cheap; a traversal that yields per step pays one round per step. That is the honest shape of the feature: it makes host round-trips expressible inside a scheme, with the same parcels and access declarations as every other node, so they can later be replaced by a device-side handler (YieldPoint::node) without touching the script.

Where it fits

Yielding scripts are the Fondaco petition mechanism made concrete on today's GPUs. The mailbox is the runtime's, the arena is the runtime's, and the only imperative host code is the handler body — the scheme stays declarative and the parcels stay in trust. See Design Thesis for the model this implements and CPU Dispatches for the simpler host-visible node it builds on.

Slang in One Source

Goldy uses Slang as its single shader language across all backends. You write one .slang file and Goldy compiles it to the native format for whichever GPU API is in use — no manual HLSL/GLSL/MSL translation, no per-backend shader files.

Compilation Targets

BackendTarget FormatAPI RequirementStatus
VulkanSPIR-VVulkan 1.4+Shipped
DirectX 12DXILWindows 10+Shipped
MetalMetal IRMetal Tier 2+ (Argument Buffers)Shipped
CUDAPTXNVIDIA CUDAIn progress
WebGPUWGSLWebGPU (via wgpu)In progress

Slang compiles through its native slang.dll / libslang.dylib — the same compiler used by NVIDIA, Khronos, and major game engines. Goldy links it directly; there is no intermediate translation step.

Why Slang

  • One source: Vertex, fragment, and compute shaders all live in a single .slang file. No preprocessor gymnastics to target different backends.
  • HLSL-compatible syntax: If you know HLSL, you already know Slang. Standard types (float4, uint3, Texture2D), standard intrinsics (mul, lerp, smoothstep), standard semantics (SV_Position, SV_Target).
  • Modern language features: Modules (import), generics, interfaces, operator overloading, and automatic differentiation — features that HLSL and GLSL lack.
  • Khronos governance: Long-term stability under open-source stewardship.

Cross-Backend Matrix Layout Consistency

Slang normalizes matrix layout across all backends. HLSL defaults to column-major storage, GLSL to column-major, and Metal to column-major — but the conventions for how mul(matrix, vector) is interpreted differ. Slang's compilation ensures that a float4x4 in your shader has identical memory layout and multiplication semantics whether it compiles to SPIR-V, DXIL, or Metal IR.

This means your Rust-side #[repr(C)] matrix types can use the same byte layout regardless of which backend the application runs on.

Shader Module Creation

Basic Compilation

ShaderModule::from_slang() compiles a Slang source string into GPU bytecode:

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang(&device, r#"
    import goldy_exp;

    [goldy_compute]
    [numthreads(64, 1, 1)]
    void cs_main(Scattered<float> data, ThreadId id) {
        data[id.x] = data[id.x] * 2.0;
    }
"#)?;
}

The goldy_exp library is pre-registered on every device — import goldy_exp works without any setup.

Additional Search Paths

ShaderModule::from_slang_with_paths() adds filesystem directories to the Slang module search path:

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang_with_paths(
    &device,
    source,
    &["my_project/shaders"],
)?;
}

Preprocessor Defines

ShaderModule::from_slang_with_paths_and_defines() passes preprocessor defines for shader variants:

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang_with_paths_and_defines(
    &device,
    source,
    &[],
    &[("msaa", "1"), ("SAMPLE_COUNT", "4")],
)?;
}

Use this for facts that are known at module creation and stay true for the life of the pipeline — a compile-time SAMPLE_COUNT, a debug dump. Do not grow a combinatorial define matrix for scene-mode flags that might hold still for a hundred frames and then change. Those belong on the dispatch as with_param scalars: the runtime's specialization predictor will bake a stable word into a _GOLDY_SPEC_* define on its own, as a full recompile of one predicted variant, and demote the moment the word moves. Author-supplied defines and runtime baking are the same compiler mechanism; they are not the same product surface.

Full Options

ShaderModule::from_slang_with_options() provides complete control — search paths, defines, optimization level, and layout validation checks:

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang_with_options(
    &device,
    source,
    &["shaders/"],
    &[("DEBUG", "1")],
    OptimizationLevel::Default,
    &[TimeUniforms::LAYOUT_CHECK],
)?;
}

Built-in Shader Modules

Goldy ships a few complete shaders as Rust string constants in goldy::shader::builtins:

ConstantDescription
VERTEX_COLOR_2D2D vertex+fragment shader with per-vertex color
SOLID_COLORSolid color fragment shader with a uniform

These are self-contained (no import needed) and useful for bootstrapping:

#![allow(unused)]
fn main() {
use goldy::shader::builtins;

let shader = ShaderModule::from_slang(&device, builtins::VERTEX_COLOR_2D)?;
}

Shader Libraries

Shader libraries are reusable Slang modules registered with a Runtime. Once registered, any shader compiled on that device can import the library.

The Built-in goldy_exp Library

Every device comes with goldy_exp pre-registered. It provides:

  • Resource type aliases (Scattered<T>, BufRO<T>, Interpolated<T>, etc.)
  • System-value wrappers (ThreadId, VertexId, InstanceId, etc.)
  • Vertex formats (FullscreenVarying, ColoredVarying, etc.)
  • Math utilities (hash(), center_uv(), smootherstep(), etc.)
  • Color utilities (rainbow(), palette(), hsv_to_rgb(), etc.)
  • Procedural geometry (quad_position(), billboard_position(), etc.)

Registering Custom Libraries

#![allow(unused)]
fn main() {
use goldy::ShaderLibrary;

device.register_library(ShaderLibrary::from_source("myutils", r#"
    module myutils;
    public float3 my_effect(float t) { return float3(t, t * 0.5, 1.0 - t); }
"#))?;
}

Now any shader can import myutils:

import myutils;

[goldy_fragment]
float4 fs_main(FullscreenVarying input) : SV_Target {
    return float4(my_effect(input.uv.x), 1.0);
}

Multi-Module Libraries

For larger libraries with internal sub-modules:

#![allow(unused)]
fn main() {
let lib = ShaderLibrary::from_embedded("effects", &[
    ("effects", r#"
        module effects;
        __include "effects/blur";
        __include "effects/bloom";
    "#),
    ("effects/blur", r#"
        implementing effects;
        public float4 gaussian_blur(Texture2D<float4> tex, SamplerState s, float2 uv) { ... }
    "#),
    ("effects/bloom", r#"
        implementing effects;
        public float4 bloom(Texture2D<float4> tex, SamplerState s, float2 uv, float threshold) { ... }
    "#),
]);

device.register_library(lib)?;
}

Loading from the Filesystem

#![allow(unused)]
fn main() {
let lib = ShaderLibrary::from_directory("effects", Path::new("shaders/effects/"))?;
device.register_library(lib)?;
}

Library Management

#![allow(unused)]
fn main() {
device.has_library("goldy_exp");       // true — always registered
device.list_libraries();               // ["goldy_exp", "myutils", ...]
device.unregister_library("myutils");  // remove a custom library
}

Layout Validation

When Rust structs are passed to shaders as uniform data (e.g. via Broadcast), the memory layout must match exactly. Goldy can validate this at compile time using Slang reflection.

Setup

  1. Derive LayoutCheckable on your Rust struct:
#![allow(unused)]
fn main() {
#[derive(LayoutCheckable)]
#[repr(C)]
struct TimeUniforms {
    time: f32,
    delta_time: f32,
    frame: u32,
    _pad: u32,
}
}
  1. Pass the layout check to shader compilation:
#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang_with_options(
    &device,
    source,
    &[],
    &[],
    OptimizationLevel::Default,
    &[TimeUniforms::LAYOUT_CHECK],
)?;
}
  1. Enable validation via environment variable:
GOLDY_VALIDATE_LAYOUTS=1 cargo run
# or
GOLDY_VALIDATION=layout cargo run
# or enable everything:
GOLDY_VALIDATION=all cargo run

What Gets Validated

  • Field offsets: Each field's byte offset in the Rust struct is compared against the Slang reflection data.
  • Struct size: Total size must match.
  • Buffer element stride: At dispatch time, the buffer's recorded element stride is checked against what the shader expects.

Validation is zero-cost when disabled — the checks are skipped entirely, not compiled out. The environment variable is read at runtime so it can be toggled without recompiling.

GOLDY_VALIDATION

The GOLDY_VALIDATION environment variable controls multiple validation categories:

ValueLayout ChecksGPU API Validation
layoutYesNo
apiNoYes
layout,apiYesYes
allYesYes
1 / true / yesNoYes

GOLDY_VALIDATE_LAYOUTS=1 is a standalone toggle that enables layout checks regardless of GOLDY_VALIDATION.

Settlement

Goldy makes GPU completion observable as settlement of concrete objects — submissions, parcels, and exchange claims — not as raw timeline numbers.

Internal clearing still uses a monotonic device clock. That clock is crate-private. Clients wait for work or resources to settle.

Submission settlement

Every successful Scheme::submit returns a Submission:

#![allow(unused)]
fn main() {
let submission = scheme.submit()?;

if !submission.is_settled() {
    submission.wait_until_settled()?;
}
}

Bounded wait:

#![allow(unused)]
fn main() {
let done = submission.wait_until_settled_timeout(1000)?; // milliseconds
if !done {
    // GPU has not finished yet
}
}

The submission owns the context it was submitted on; callers do not pass a Context to wait.

Parcel and resource settlement

Before reusing or dropping a resource that may still be referenced by in-flight GPU work:

#![allow(unused)]
fn main() {
if !parcel.is_settled() {
    parcel.wait_until_settled()?;
}
}

The same methods exist on Buffer and Texture.

Direct host writes on CPU-writable buffers require the buffer to be settled (or never GPU-referenced). Prefer MemoryExchange deposits for uploads.

Exchange claims (unchanged)

Surface and memory exchanges still settle occurrences via consume/discard. Rust surface present sugar is (&mut submission >> &transaction).take()?; the &mut borrow is required by operator semantics and leaves other claims untouched.

#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;

// Present
(&mut submission >> &transaction).take()?;

// Host claim — wait + mapped pointer or staged copy
let view = (&mut submission >> &parcel).take::<u32>()?;
}

A live linear claim is unsettled until consume or discard. Dropping an unsettled claim discards it.

Memory deposits follow the same grammar internally (Transaction → claim at submit → consume at the copy dispatch) but the program never authors the claim. bind_deposit records copy topology; (&deposit << &data)? (or DepositTransaction::write) prepares the occurrence for this submission; submit claims it; graph execution consumes it. Exchange staging is retired locally and is not a parcel-ledger entry. Destination RAW/WAR ordering remains enforced.

Multi-frame pipelining

For production renderers, use FrameOrchestrator. It bounds CPU/GPU depth using submissions — not raw epochs:

#![allow(unused)]
fn main() {
let mut orch = FrameOrchestrator::new(&ctx, 3);

loop {
    let handle = orch.begin_frame()?;
    let submission = scheme.submit()?;
    orch.end_frame_standalone(handle, &submission)?;
}

orch.drain_all()?;
}

How this differs from fence-based APIs

Traditional GPU APIs expose fence objects or timeline counters to the application. Goldy keeps those as runtime clearing instruments (finance analogy: sequence numbers in a clearinghouse). Application code holds receipts (Submission) and parcels (Parcel) and asks when those are settled.

Fence / timeline counterSettlement
QueryPoll a fence or compare u64obj.is_settled()
WaitWait on fence / wait_until(tv)obj.wait_until_settled()
IdentityOpaque fence or epoch numberConcrete submission or parcel
PortabilityTied to native timeline primitivesBackend may use fences, events, or onSubmittedWorkDone

Resource lifetime

Dropping a Buffer or Texture may be deferred internally until GPU work that referenced it has retired. Prefer settling before dropping when you need deterministic reclaim timing (for example Metal heap-sensitive resize paths).

Tensor Algebra

Goldy's tensor feature is a batteries-included dense tensor algebra over parcels. It is not an ML framework: there is no autograd, no nn.Module, no optimizer, and no Llama-specific operators. A future neural-network library can compete with torch.nn by building on this layer.

flowchart TD
  App["Application or model"] --> Tensor["Goldy tensor layer"]
  Tensor --> Scheme["Goldy Scheme and parcels"]
  Scheme --> Backend["CUDA, Metal, Vulkan, DX12, WebGPU, CPU"]
  NN["Future NN library"] --> Tensor

Enable it with the tensor Cargo feature (on by default; independent of graphics):

goldy = { version = "0.3", features = ["tensor"] }
# CUDA compute-only, no graphics:
# goldy = { version = "0.3", default-features = false, features = ["cuda", "tensor"] }

What the layer owns

  • Concrete ranks 0..=4, dtypes F32 / U32 / I32
  • Packed row-major storage plus positive/zero-stride views
  • Checked narrow, reshape, permute, transpose, broadcast_to
  • Shape inference, output allocation, and recording into the same Scheme
  • Portable elementwise / reduction / gather / scatter kernels plus a tensor front-end for semantic MatMul

What it does not own

Autograd, parameters, optimizers, model formats, RMSNorm, RoPE, attention, activations as NN layers, KV-cache policy, tokenization, symbolic shapes, or negative strides.

Low-level Scheme, Buffer, Parcel, and MatMulView APIs remain escape hatches.

Views are lenses

A [Tensor] is a buffer parcel plus layout metadata. A [TensorView] never mints a new ownership identity. Binding a view:

  1. Claims a conservative byte envelope (BufferRange, or the parent buffer) so overlapping aliases stay visible to Goldy's hazard analysis.
  2. Uses the parent bindless slot so shaders see the whole buffer plus element offsets.

Read-only broadcast views (zero stride on an expanded axis) share that envelope; they are not independent identities. Write layouts require a positive stride on every axis with shape > 1.

Host updates and observations still go through MemoryExchange.

Recording

[TensorKernels] prepares portable kernels. Layout parcels intern onto the Scheme at record time, so the kernels object only needs to live while you are recording. [TensorRecorder] borrows the kernels and a mutable Scheme:

use goldy::{Instance, RequestAdapterOptions, RuntimeDescriptor, Scheme, Tensor, TensorKernels, TensorShape};
fn main() -> Result<(), goldy::GoldyError> {
let runtime = Instance::new().unwrap().request_adapter(&Default::default()).unwrap().request_runtime(&Default::default()).unwrap();
let ctx = runtime.create_context().unwrap();
let kernels = TensorKernels::new(&runtime)?;
let a = Tensor::from_f32(&runtime, TensorShape::vector(4), &[1.0, 2.0, 3.0, 4.0])?;
let b = Tensor::from_f32(&runtime, TensorShape::vector(1), &[10.0])?;
let mut scheme = Scheme::new(&ctx);
let c = kernels.recorder(&mut scheme).add("add", a.view(), b.view())?;
let _ = c;
Ok(())
}

Allocating methods (add, matmul, sum, …) create packed outputs. _into forms (add_into, matmul_into, cast_into, fill) use caller storage, including in-place work when the write layout is legal.

Custom #[goldy::compute] kernels can take gpu::Tensor<T> parameters and index them as logical views, or bind Tensor / TensorView like any other parcel for physical indexing:

// Logical: view[i] applies offset/shape/strides. Layouts live on the scheme.
kernel.record(&mut scheme, "rope", q_view, k_layer, &step, theta)?
    .over_tensor(&q_view);

// Physical escape hatch: buf[i] is a parent-buffer element index.
kernel.record(&mut scheme, "double", &data.view(), n).over_tensor(&data.view());

Kernel parameters may declare a shape contract that record checks before GraphIR insertion. The list fixes rank; repeated names must match; _ is unconstrained; integer literals are exact extents. Unannotated tensors stay any-shape. See Rust compute kernels.

fn rmsnorm(
    #[tensor(shape = [dim])] x: gpu::Tensor<f32>,
    #[tensor(shape = [dim])] weight: gpu::Tensor<f32>,
    #[tensor(shape = [dim])] out: gpu::TensorWrite<f32>,
) { /* ... */ }
fn rope(
    #[tensor(shape = [q_heads, head])] q: gpu::TensorMut<f32>,
    #[tensor(shape = [seq, kv_heads, head])] k: gpu::TensorMut<f32>,
    step: &[DecodeStep],
    theta: f32,
) {
    let k_base = pos * k.dim(1) * k.dim(2);
    k[k_base + i] = ...;
}
// host: pass layout.embedding(weights)? and layer_cache(key_cache, layer)?

Packed checkpoint pointer walking stays in the ingestion crate. It is the single boundary that translates foreign offsets into validated TensorViews.

Operations

FamilyOpsNotes
Constructionzeros, from_f32 / from_u32 / from_i32, fill, copy, contiguous, castPer-op dtype support
Unaryneg, abs, exp, log, sqrt, reciprocalF32
Binaryadd / sub / mul / div / min / max and *_scalarF32, NumPy-style broadcasting
Reductionssum, max_reduce, min_reduce, meanOne axis; squeezed by default
Gather / scattergather; scatter with [ScatterMode]Modes are explicit: UniqueWrite, Add, Min, Max
Matmulrank-1/2 and batched rank-3Rank-2 packed cases lower to Scheme::matmul (cuBLAS / MPS / stdlib)
Softmaxcomposition of max / sub / exp / sum / divConvenience only; not fused attention

Scatter Add / Min / Max are defined (a single thread walks colliding indices). They are not idempotent on scheme replay if the destination is the accumulator; unique-write and pure functions of the inputs are.

Bindings

The tensor Cargo feature is passed through to goldy-ffi, goldy-ffi-client, and goldy-py. C (goldy.h) and C++ (goldy.hpp) expose acquire / add / matmul / fill. Python copies NumPy arrays into tensors on acquire and reads results back through MemoryExchange on tensor.parcel() — there is no arbitrary zero-copy host view of GPU storage.

llama3.goldy

llama3.goldy is the proving consumer: activations and checkpoint weights are tensors. Static GEMVs and residuals go through Ammon's TensorKernels (Goldy semantic matmul / portable add). Custom RMSNorm / RoPE / attention / SwiGLU kernels bind tensor views with rank and symbolic-extent contracts; the host reshapes Q, KV cache, and attention scores at recording boundaries. Dynamic decode state stays a DecodeStep deposit.

Matrix Multiply

Scheme::matmul records a semantic GEMM/GEMV. The task graph schedules it from buffer bindings like any other node. The backend chooses an implementation on the first submit and retains that plan.

The tensor front end (goldy/tensor, on by default) derives m/n/k, transpose flags, offsets, and leading dimensions from checked TensorViews and records the same node. MatMulView remains the low-level escape hatch.

#![allow(unused)]
fn main() {
scheme
    .matmul("q_projection", MatMulDesc::gemv(dim, dim))
    .a(&weights, MatMulView::offset(wq))
    .b(&xb, MatMulView::packed())
    .out(&q, MatMulView::packed())
    .record();
}

The contract is row-major C[m, n] = alpha * op(A)[m, k] @ op(B)[k, n] + beta * C. The first slice is FP32 with alpha = 1, beta = 0 (the stdlib fallback requires those epilogue values; native libraries honor other alpha/beta).

Implementation choice

BackendDefaultOverride
CUDAcuBLAS (cublasSgemv when n = 1, otherwise cublasSgemm)GOLDY_MATMUL=fallback
MetalMetal Performance ShadersGOLDY_MATMUL=fallback
Vulkan, DX12, WebGPU, CPUGoldy stdlib kernel—

There is no public prepare(). The stdlib pipeline is compiled on first submit when the backend has no native library (or when fallback is forced). Subsequent clean submits reuse the realized command list / CUDA graph.

Custom leading dimensions are honored by native libraries. The stdlib kernel requires packed row-major storage (lda/ldb/ldc derived from m/n/k and the transpose flags).

Pipelined Frames

Goldy's FrameOrchestrator manages CPU/GPU frame pacing: an in-flight ring, a depth cap, and present-path settlement patching. It does not own GPU bytes or run cleanup callbacks — recycle lives in TransientPool and Runtime.

The problem it solves

Every pipelined renderer needs the same bookkeeping:

  1. A ring of in-flight frame receipts.
  2. A pipeline-depth cap — block the CPU when the ring is full to prevent unbounded memory growth.
  3. Deferred retirement of slots when their submissions settle.
  4. Present-path patching — stamp the most recent slot after Claim::consume via note_presented.

Without shared infrastructure, every consumer reimplements this independently. FrameOrchestrator centralizes the pacing half of it.

Core API

#![allow(unused)]
fn main() {
use goldy::{FrameOrchestrator, FrameHandle};

// max_depth: how many frames may be in-flight before begin_frame blocks
let mut orch = FrameOrchestrator::new(&ctx, 3);
}

Standalone (headless / render-to-texture) path

#![allow(unused)]
fn main() {
loop {
    // 1. Open a new frame slot; drains completed older slots.
    //    Blocks if max_depth frames are already in flight.
    let handle = orch.begin_frame()?;

    // 2. Submit retained scheme work (recorded earlier or this frame).
    let submission = scheme.submit()?;

    // 3. Register the slot from the submission.
    orch.end_frame_standalone(handle, &submission)?;
}
}

Present-on-scheme (swapchain) path

#![allow(unused)]
fn main() {
loop {
    let handle = orch.begin_frame()?;

    let mut submission = scheme.submit()?;
    (&mut submission >> &present).take()?;

    orch.end_frame_for_present(handle, &submission)?;
    orch.note_presented(&submission);
}
}

Externally ordered path

When scheme submit sidecars / present easement already enforce cross-frame ordering, close with end_frame_externally_ordered so no ring slot is created and the next begin_frame does not wait on a coarse frame timeline.

Mid-frame submit boundaries

Split a frame into multiple scheme submissions so the GPU can begin earlier phases while the CPU records later ones:

#![allow(unused)]
fn main() {
let handle = orch.begin_frame()?;

// Coarse phase
let _coarse = coarse_scheme.submit()?;

// Fine phase — GPU executes coarse while CPU records/submits this
let fine = fine_scheme.submit()?;

orch.end_frame_standalone(handle, &fine)?;
}

Each Scheme::submit creates a real command-buffer boundary on all backends. Because Metal (and Vulkan/DX12) execute command buffers on the same queue in submission order, the fine submission automatically waits for the coarse one — no explicit fence is required.

CPU/GPU overlap

FrameOrchestrator enables two distinct layers of CPU/GPU overlap:

Frame-level — begin_frame drains completed slots without blocking when under the depth cap, so the CPU can start recording frame N+1 while the GPU executes frame N. The depth cap (max_depth) prevents the CPU from running too far ahead.

Intra-frame — multiple scheme.submit() calls in one frame split the command stream into multiple GPU submissions. The GPU starts executing the first submission before the CPU finishes the last one.

Inspecting orchestrator state

#![allow(unused)]
fn main() {
orch.pending_frames();   // slots currently in the ring
orch.max_depth();        // cap configured at construction
orch.has_open_frame();   // true between begin_frame and end_frame_*
}

Under allocation pressure, orch.wait_for_progress() blocks on the oldest ring slot (or flushes deferred deletions when the ring is empty).

Design notes

Present path settlement is always deferred

On the swapchain path the final scanout settlement may arrive only after Claim::consume. The orchestrator holds the slot unset until note_presented arrives.

Relationship to resource recycling

FrameOrchestrator owns the frame-slot ring only. Transient buffer/texture recycling lives in the per-context TransientPool (leases via acquire_transient_* / return_transient_*). They are independent: the orchestrator does not call into the pool, and clients must not hang byte reclaim on orchestrator callbacks (there are none).

Compute to Surface

Compute-to-surface lets a compute shader write directly to a swapchain drawable, bypassing the rasterization pipeline entirely. There is no RenderPipeline, no vertex buffers, and no raster pass — just a compute dispatch that fills pixels.

When to use compute-to-surface

Use compute-to-surface when your rendering is naturally a per-pixel computation rather than geometry rasterization:

  • Fullscreen image effects (plasma, fractals, ray marching)
  • GPU-driven 2D renderers where the compute shader owns the output layout
  • Post-processing that doesn't need triangle rasterization
  • Prototyping visual effects without setting up a render pipeline

Use traditional rendering when you need the rasterization pipeline's features: triangle assembly, depth testing, MSAA, alpha blending, or vertex/fragment shader stages.

Surface exchange

Create a SurfaceExchange and call bind_destination to register direct compute-to-present in the scheme:

#![allow(unused)]
fn main() {
let surface = SurfaceExchange::new_with_depth(&ctx, &window, 3, SurfaceConfig::default())?;
let (lease, present) = surface.bind_destination(&mut scheme)?;
}

Bind the returned lease in a compute node with with_present(&lease). Goldy handles barrier insertion between compute writes and the presentation engine.

On CUDA+DX12, present scratch is a depth-3 imported staging ring (not the DXGI backbuffer). Compute does not reuse frame N's scratch in frame N+1; that extra memory is the interop tradeoff until CUDA/DX12 synchronization APIs improve.

Building the scheme

Record a retained scheme with a compute node that writes to the present lease:

#![allow(unused)]
fn main() {
let wg_x = width.div_ceil(8);
let wg_y = height.div_ceil(8);

let mut scheme = Scheme::new(&ctx);
let (lease, present) = surface.bind_destination(&mut scheme)?;
scheme
    .node("compute", &compute_pipeline)
    .with_parcel(&uniform_buffer, NodeAccess::Read)
    .with_present(&lease)
    .dispatch(wg_x, wg_y, 1);
}

Submitting and presenting

Each frame, submit the scheme and consume the surface claim:

#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

submit resolves transient resources, compiles the scheme into a command stream, and submits to the GPU. Presentation happens when you call (&mut submission >> &present).take()? — the compute shader has already written the pixels. The explicit &mut borrow is required by operator semantics and leaves other claims on the submission untouched.

The compute shader

The shader receives the output texture as a DirectSpatial<float4> — a read-write 2D texture accessed by integer coordinates:

import goldy_exp;

struct Uniforms {
    uint width;
    uint height;
    float time;
};

[goldy_compute]
[numthreads(8, 8, 1)]
void cs_main(BufRO<Uniforms> uniforms_buf, DirectSpatial<float4> output, ThreadId tid) {
    Uniforms u = uniforms_buf[0];

    if (tid.x >= u.width || tid.y >= u.height)
        return;

    float2 uv = float2(float(tid.x) / float(u.width),
                       float(tid.y) / float(u.height));

    // Compute pixel color...
    float3 col = my_color_function(uv, u.time);
    output[tid.xy] = float4(col, 1.0);
}

The [numthreads(8, 8, 1)] workgroup size maps naturally to 2D image tiles. Dispatch enough workgroups to cover the full resolution:

#![allow(unused)]
fn main() {
let wg_x = width.div_ceil(8);
let wg_y = height.div_ceil(8);
}

Guard against out-of-bounds writes in the shader when the resolution isn't a multiple of the workgroup size.

Full example sketch

#![allow(unused)]
fn main() {
use goldy::{
    BufferKind, ComputePipeline, RuntimeDescriptor, Instance, MemoryExchange, NodeAccess, PresentMode,
    RequestAdapterOptions, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange,
};

let instance = Instance::new()?;
let device = instance
    .request_adapter(&RequestAdapterOptions::default())?
    .request_runtime(&RuntimeDescriptor::default())?;
let ctx = device.create_context()?;

let surface = SurfaceExchange::new_with_config(
    &ctx,
    &window,
    SurfaceConfig {
        present_mode: PresentMode::Fifo,
        depth_format: None,
    },
)?;

let shader = ShaderModule::from_slang(&device, COMPUTE_SHADER)?;
let compute_pipeline = ComputePipeline::new(&device, &shader)?;

let uniform_buffer = device.acquire_buffer_with_data(
    &[Uniforms { width, height, time: 0.0 }],
    BufferKind::Scattered,
)?;

let mut scheme = Scheme::new(&ctx);
let (lease, present) = surface.bind_destination(&mut scheme)?;
scheme
    .node("compute", &compute_pipeline)
    .with_parcel(&uniform_buffer, NodeAccess::Read)
    .with_present(&lease)
    .dispatch(width.div_ceil(8), height.div_ceil(8), 1);

// --- Render loop ---
let mut upload = Scheme::new(&ctx);
let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
    &mut upload,
    goldy::DepositTarget::buffer_elements::<Uniforms>(&uniform_buffer, 1),
)?;
(&uniform_deposit << &Uniforms { width, height, time: elapsed })?;
upload.submit()?;

let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

See examples/compute_to_surface.rs for the complete winit application.

Pipelines

Pipelines combine compiled shaders with fixed-function rendering state. Goldy provides RenderPipeline for raster, MeshPipeline for mesh shading, and ComputePipeline for compute.

A graphics pipeline is a typed connection:

vertex input → vertex/mesh stage → interpolated payload → fragment stage

Think of it as RasterPipeline<VertexIn, Varying, Color> — without repeating those types in Rust. Payloads stay shader-owned. Goldy reflects each stage, links them structurally by semantic, and merges per-stage virtual-main resources into one named contract.

Three things a draw needs

ConceptWho owns itExample
Pipeline payloadShader structs with semantics (SV_Position, TEXCOORD0)Varying returned by the vertex/mesh stage and consumed by the fragment stage
Runtime resource bindingsNamed virtual-main parameters, bound at record timeShaderBinding::read("scene", &scene)
Invocation builtinsSystem-value wrappersVertexId, ThreadId, IsFrontFace

RenderPipeline::interface() exposes the reflected vertex input, payload links, fragment outputs, and merged resource contract for diagnostics.

Render Pipelines

A RenderPipeline pairs vertex and fragment shaders with raster state: vertex input layout, primitive assembly, depth testing, and the output format.

Creating a Render Pipeline

The builder reflects and links stages automatically. Existing RenderPipeline::new delegates to the same linker.

#![allow(unused)]
fn main() {
use goldy::{RenderPipeline, ShaderBinding, ShaderModule, Vertex2D};

let vs = ShaderModule::from_slang(&device, include_str!("shaders/tri.vs.slang"))?;
let fs = ShaderModule::from_slang(&device, include_str!("shaders/tri.fs.slang"))?;

let pipeline = RenderPipeline::builder(&device)
    .vertex(&vs)
    .fragment(&fs)
    .vertex_layout(Vertex2D::layout())
    .topology(PrimitiveTopology::TriangleList)
    .target_format(surface.format())
    .build()?;
}

Each stage may declare only the resources it uses. Shared names (scene on both VS and FS) merge into one slot; unique names append after the fragment list (raster) or the mesh list (mesh pipelines).

Named draw bindings

#![allow(unused)]
fn main() {
pass.with_shader_bindings(&[
    ShaderBinding::read("scene", &scene),
    ShaderBinding::read("albedo", &albedo),
    ShaderBinding::sampler("nearest", &sampler),
]);
pass.set_pipeline(&pipeline);
}

Extra names are allowed so one pass-level set can serve several pipeline switches. Missing required names and wrong resource categories fail with an error that names the pipeline parameter.

Positional with_shader_resources remains valid: slots follow the merged contract order (fragment-first for raster, mesh-first for mesh).

RenderPipelineDesc

RenderPipeline::new(&device, &vs, &fs, &desc) still works. The descriptor is the same raster state the builder sets:

#![allow(unused)]
fn main() {
pub struct RenderPipelineDesc {
    pub vertex_layout: VertexBufferLayout,
    pub topology: PrimitiveTopology,
    pub target_format: TextureFormat,
    pub depth_stencil: Option<DepthStencilState>,
}
}
FieldPurposeDefault
vertex_layoutDescribes vertex buffer stride and attributesEmpty (no vertex input)
topologyHow vertices are assembled into primitivesTriangleList
target_formatPixel format of the render target — must match surface.format() or the format passed to RenderTarget::new()Rgba8Unorm
depth_stencilDepth/stencil test configuration, or None to disableNone

The default descriptor is valid for fullscreen passes that generate geometry from SV_VertexID and render to an Rgba8Unorm target without depth testing.

Format Matching

The pipeline's target_format must match the render target it will draw into. Mismatched formats produce backend errors or undefined output.

#![allow(unused)]
fn main() {
let desc = RenderPipelineDesc {
    target_format: surface.format(),
    ..Default::default()
};
}

Vertex Buffer Layouts

A VertexBufferLayout tells the pipeline how to interpret vertex buffer memory. For passes that do not use vertex buffers (fullscreen triangles, quad instancing), the default empty layout is correct.

For typed vertex input, use the from_formats builder or a built-in type's layout() method. See Vertex Types and Layouts for details.

#![allow(unused)]
fn main() {
let layout = VertexBufferLayout::from_formats::<MyVertex>(&[
    VertexFormat::Float32x3, // position
    VertexFormat::Float32x2, // uv
]);
}

Primitive Topology

Controls how the vertex stream is assembled into geometric primitives:

#![allow(unused)]
fn main() {
pub enum PrimitiveTopology {
    PointList,
    LineList,
    LineStrip,
    TriangleList,   // default
    TriangleStrip,
}
}
PointList:     •  •  •  •
LineList:      •——•  •——•
LineStrip:     •——•——•——•
TriangleList:  △  △
TriangleStrip: △▽△▽

Depth/Stencil State

Enable depth testing by setting depth_stencil. The surface or render target must have been created with a matching depth format.

#![allow(unused)]
fn main() {
use goldy::{DepthStencilState, DepthFormat, CompareFunction};

let pipeline = RenderPipeline::new(&device, &vs, &fs, &RenderPipelineDesc {
    vertex_layout: Vertex2D::layout(),
    target_format: surface.format(),
    topology: PrimitiveTopology::TriangleList,
    depth_stencil: Some(DepthStencilState {
        format: DepthFormat::Depth32Float,
        depth_write_enabled: true,
        depth_compare: CompareFunction::Less,
    }),
})?;
}

DepthStencilState fields:

FieldPurposeDefault
formatDepth texture format (Depth16Unorm, Depth24Plus, Depth32Float, etc.)Depth24Plus
depth_write_enabledWhether fragments write to the depth buffertrue
depth_compareComparison function — Less, LessEqual, Greater, Always, etc.Less

Available depth formats:

FormatBitsStencil
Depth16Unorm16-bitNo
Depth24Plus24-bit (may use 32 internally)No
Depth24PlusStencil824-bit + 8-bit stencilYes
Depth32Float32-bit floatNo
Depth32FloatStencil832-bit float + 8-bit stencilYes

For reverse-Z rendering, use CompareFunction::Greater and clear depth to 0.0.

Compute Pipelines

ComputePipeline wraps a single compute shader. See Your First Compute Shader for the full compute API.

#![allow(unused)]
fn main() {
use goldy::{ComputePipeline, ShaderModule};

let cs = ShaderModule::from_slang(&device, include_str!("shaders/sim.cs.slang"))?;
let pipeline = ComputePipeline::new(&device, &cs)?;
}

Why Goldy Has Fewer Pipelines

Pipeline State Object (PSO) explosion is one of the biggest pain points in modern graphics. Engines routinely manage thousands of pipeline permutations and ship massive shader caches. Goldy eliminates most combinatorial dimensions:

DimensionTraditional Vulkan/DX12Goldy
Render pass compatibilityN render passes × M subpassesEliminated — dynamic rendering
Descriptor set layoutsPer-material layout permutationsOne global bindless layout
Pipeline layoutsPer-materialOne shared layout
Viewport / scissorBaked into PSODynamic state
Vertex formatBakedBaked (unavoidable)
Target formatBakedBaked (unavoidable)

RenderPipelineDesc has exactly four fields. The permutation space is vertex_layouts × topologies × target_formats × depth_configs — deliberately small.

Compute is the one place Goldy will compile an extra program after startup: a retained dispatch whose with_param scalars hold still is moved onto a baked variant of its shader. That is at most one specialized pipeline per dispatch site, not a permutation of feature flags, and it is an implementation detail of scheme submit rather than something ComputePipeline::new enumerates. See Shader Specialization Prediction.

Performance

Pipelines are expensive to create (shader compilation, PSO allocation) but cheap to bind during rendering. Create them once at startup and reuse across frames.

#![allow(unused)]
fn main() {
struct Renderer {
    scene_pipeline: RenderPipeline,
    ui_pipeline: RenderPipeline,
    wireframe_pipeline: RenderPipeline,
}

impl Renderer {
    fn new(device: &Runtime, surface: &Surface) -> Result<Self> {
        // Create all pipelines upfront
        Ok(Self {
            scene_pipeline: create_scene_pipeline(device, surface.format())?,
            ui_pipeline: create_ui_pipeline(device, surface.format())?,
            wireframe_pipeline: create_wireframe_pipeline(device, surface.format())?,
        })
    }
}
}

Mesh pipelines

MeshPipeline replaces the vertex stage with a mesh shader ([goldy_mesh] / mesh_main) and a fragment shader. Amplification / task shaders are optional ([goldy_amplification] / amp_main) when RuntimeCapabilities::amplification_shaders is set.

Create the pipeline only when device.capabilities().mesh_shaders is true. Record with set_mesh_pipeline and dispatch_mesh instead of draw. Vulkan, DX12, and Metal implement this.

#![allow(unused)]
fn main() {
let pipeline = MeshPipeline::builder(&device)
    .mesh(&shader)
    .fragment(&shader)
    .target_format(rt_format)
    .build()?;

let mut pass = scheme.render_pass("mesh", &rt, TargetLoad::Discard);
pass.set_mesh_pipeline(&pipeline);
pass.dispatch_mesh(1, 1, 1);
pass.finish();
}

Render Pass Nodes

Goldy has no command buffers and no command lists. A graphics draw is a render pass node inside a Scheme — the same retained dependency graph that holds compute dispatches, copies, and present nodes. scheme.render_pass(...) returns a builder; what you call on that builder is recorded into the node, not executed immediately. Nothing touches the GPU until scheme.submit().

This matters for how you think about the API: there is no "encoder" you open and close per frame. You build the graph once — typically at init and on resize — and resubmit it every frame. See Settlement for what happens after submit(), and Pipelines for how RenderPipeline fits into a pass.

finish() is the node's terminator — the same role dispatch(x, y, z) plays for compute — because a raster node is a list of draws against one target, not one launch. It is not vkCmdEndRendering; recording is pure IR. Drop does not commit the node.

Why one node for many draws

The Fondaco machine allows each draw to be its own dispatch. Goldy clusters them because 2026 graphics APIs only let you fence at the framebuffer epoch (begin/end rendering, Metal render encoder), not at each draw. Draws that share a color/depth target need that epoch so the tile stays on-chip; vertex buffers used by only the first draw still cannot be recycled until the whole pass ends. That is a substrate limit, not a scheme-theory limit. Full argument: Render Passes and Schemes.

Recording a Render Pass Node

scheme.render_pass(label, target, color_load) opens a builder bound to one leased render target:

#![allow(unused)]
fn main() {
use goldy::{Color, NodeAccess, Scheme, TargetLoad};

let mut pass = scheme.render_pass("triangle", &scene_rt, TargetLoad::Clear(Color::CORNFLOWER_BLUE));
pass.with_parcel(&vertex_buffer, NodeAccess::Read);
pass.set_pipeline(&pipeline);
pass.set_vertex_buffer(0, &vertex_buffer);
pass.draw(0..3, 0..1);
pass.finish();
}

finish() pushes the node into the scheme's graph. The builder cannot be reused after finish() — record a new pass for the next node.

Color Load

Load behavior is a property of the pass node, not a separate clear call:

VariantEffect
TargetLoad::LoadPreserve prior color contents (the node reads the target)
TargetLoad::Clear(color)Clear to color at pass start (private-inaugural — the node owns the target outright)
TargetLoad::DiscardPrior contents are irrelevant; the pass must fully overwrite every pixel

This is a scheduling input, not cosmetic: Clear/Discard tell Goldy the pass does not depend on the target's previous contents, which affects how the runtime orders and aliases transient render targets across the scheme.

Depth

#![allow(unused)]
fn main() {
pass.clear_depth(1.0);
}

Depth clear is declared the same way — as part of the node, before drawing.

Declaring Dependencies

A render pass node participates in the scheme's dependency graph the same way a compute node does. Declare every parcel it reads or writes so Goldy can derive barriers and track parcel lifetimes:

#![allow(unused)]
fn main() {
pass.with_parcel(&vertex_buffer, NodeAccess::Read);
pass.with_parcel(&uniform_buf, NodeAccess::Read);
}

with_parcel also registers the parcel for typed bindless binding, in call order, the next time set_pipeline is called — so declare dependencies for a draw before calling set_pipeline for it.

For a Buffer you want to depend on without binding it as a shader resource (e.g. a geometry buffer accessed only through set_vertex_buffer/set_index_buffer), use with_buffer_dependency instead — it registers the dependency without claiming a bindless slot:

#![allow(unused)]
fn main() {
pass.with_buffer_dependency(&geometry, NodeAccess::Read);
}

Pipeline, Buffers, and Drawing

#![allow(unused)]
fn main() {
pass.set_pipeline(&pipeline);
pass.set_vertex_buffer(0, &vertex_buffer);
pass.set_index_buffer(&indices, IndexFormat::Uint16);

pass.draw(0..3, 0..1);              // non-indexed: vertex range, instance range
pass.draw_indexed(0..6, 0..1);      // indexed: index range, instance range
pass.draw_fullscreen();             // shorthand for draw(0..3, 0..1)

pass.set_mesh_pipeline(&mesh_pipeline);
pass.dispatch_mesh(1, 1, 1);        // mesh workgroups (Vulkan / DX12 / Metal)
}

Do not mix draw with a mesh pipeline or dispatch_mesh with a vertex pipeline — Scheme::submit returns GoldyError::Validation with a hint. Create MeshPipeline only when device.capabilities().mesh_shaders is true (see Target Hardware).

set_pipeline binds a RenderPipeline and, if any parcels were declared with with_parcel beforehand, resolves and binds their bindless handles for that pipeline's typed shader parameters. Calling set_pipeline again mid-pass starts a new binding scope for subsequent draws — declare each draw's parcels right before the set_pipeline call that will consume them.

For fullscreen or procedurally-generated geometry (no vertex buffer at all), skip set_vertex_buffer entirely and generate positions from SV_VertexID in the shader — see Vertex Types and Layouts.

Offscreen-Only (Tests, Readback)

Headless rendering — no window, no SurfaceExchange — records the same render pass node, then host-claims pixels from a copy destination:

#![allow(unused)]
fn main() {
let mut scheme = Scheme::new(&ctx);
let rt = ctx.lease_render_target(800, 600, TextureFormat::Rgba8Unorm, None)?;

let mut pass = scheme.render_pass("clear", &rt, TargetLoad::Clear(Color::RED));
pass.finish();

scheme.copy_to_texture(&rt, &readback_texture);
let mut submission = scheme.submit()?;
let pixels = (&mut submission >> &readback_texture).take::<u8>()?.to_vec();
}

Windowed Rendering

A windowed frame is the same render pass node, plus a present binding recorded once against the scheme via SurfaceExchange:

#![allow(unused)]
fn main() {
use goldy::{Color, NodeAccess, Scheme, SurfaceExchange, TargetLoad};

// Record once, at init and on resize:
let mut pass = scheme.render_pass("main", &scene_rt, TargetLoad::Clear(Color::CORNFLOWER_BLUE));
pass.with_parcel(&vertex_buffer, NodeAccess::Read);
pass.set_pipeline(&pipeline);
pass.set_vertex_buffer(0, &vertex_buffer);
pass.draw(0..3, 0..1);
pass.finish();
let present = surface.bind_render_target(&mut scheme, &scene_rt)?;

// Each frame:
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

The graph is recorded once; every frame just resubmits it and settles the present claim. See examples/triangle.rs for the full loop, including resize handling (rebuild the scheme and transaction when the surface size changes).

Compute and Graphics in One Scheme

Because render pass nodes and compute nodes live in the same graph, a hybrid frame is just multiple node(...) and render_pass(...) calls on one Scheme, submitted together:

#![allow(unused)]
fn main() {
let memory = MemoryExchange::new(&ctx);
let deposit = memory.bind_deposit(&mut scheme, goldy::DepositTarget::buffer(&staging, data.len() as u64))?;
(&deposit << data.as_slice())?;

scheme.node("sim", &compute_pipeline)
    .with_parcel(&state_buf, NodeAccess::ReadWrite)
    .dispatch(wg, 1, 1);

let mut pass = scheme.render_pass("draw", &scene_rt, TargetLoad::Discard);
pass.with_parcel(&state_buf, NodeAccess::Read);
pass.set_pipeline(&pipeline);
pass.set_vertex_buffer(0, &vertex_buffer);
pass.draw(0..3, 0..1);
pass.finish();

let present = surface.bind_render_target(&mut scheme, &scene_rt)?;
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

Goldy derives the ordering between the compute node and the render pass node from their declared parcel accesses — the simulation's write to state_buf is ordered before the pass's read, with no barrier authored by hand.

Notes

  • A render pass builder is single-use: call finish() (Rust) once recording is complete, before scheme.submit(). In the FFI bindings (C++, .NET, goldy-ffi-client), the equivalent is a RAII scope that finishes on drop or block exit.
  • A pass node is scoped to one leased render target for its lifetime — draw into a different target by opening a new render_pass(...) node.
  • Nothing in this page executes anything: recording is pure graph-building. Execution, barrier insertion, and transient aliasing all happen inside scheme.submit().
  • A draw inside the builder is not a waitable epoch. Compute that overwrites a vertex parcel, or samples the color target, is a later scheme node. Splitting the pass (TargetLoad::Load on a second node) is how you manufacture a gate; it stores and reloads the attachment.

Vertex Types and Layouts

Goldy provides built-in vertex types for common 2D rendering and a layout builder for custom vertex formats. Vertex data is described by a VertexBufferLayout that tells the pipeline how to interpret buffer memory.

Built-in Vertex Types

Vertex2D

Position + color. Use for colored primitives, particles, and debug visualization.

#![allow(unused)]
fn main() {
use goldy::{Vertex2D, Color};

let vertices = vec![
    Vertex2D::new(-0.5, -0.5, Color::RED),
    Vertex2D::new( 0.5, -0.5, Color::GREEN),
    Vertex2D::new( 0.0,  0.5, Color::BLUE),
];
}

Memory layout (24 bytes per vertex):

LocationFieldFormatOffset
0positionFloat32x20
1colorFloat32x48

Get the pipeline layout with Vertex2D::layout().

Vertex2DUv

Position + texture coordinates. Use for textured quads, sprites, and shader effects.

#![allow(unused)]
fn main() {
use goldy::Vertex2DUv;

let vertices = vec![
    Vertex2DUv::new(-1.0, -1.0, 0.0, 1.0),
    Vertex2DUv::new( 1.0, -1.0, 1.0, 1.0),
    Vertex2DUv::new( 0.0,  1.0, 0.5, 0.0),
];
}

Memory layout (16 bytes per vertex):

LocationFieldFormatOffset
0positionFloat32x20
1uvFloat32x28

Get the pipeline layout with Vertex2DUv::layout().

Using Built-in Types in Pipelines

Both types provide a layout() method that returns the correct VertexBufferLayout:

#![allow(unused)]
fn main() {
let pipeline = RenderPipeline::new(&device, &vs, &fs, &RenderPipelineDesc {
    vertex_layout: Vertex2D::layout(),
    target_format: surface.format(),
    ..Default::default()
})?;
}

Both types implement StructuredBufferElement, so they can also be stored via Runtime::acquire_buffer_with_data.

Custom Vertex Layouts

Defining a Custom Vertex

Custom vertex types must be #[repr(C)] and derive bytemuck::Pod and bytemuck::Zeroable:

#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
struct MyVertex {
    position: [f32; 3],
    normal: [f32; 3],
    uv: [f32; 2],
    color: u32,
}
}

Building a Layout with from_formats

VertexBufferLayout::from_formats::<T> infers locations (sequential from 0) and offsets (accumulated from format sizes), then validates that the total matches size_of::<T>():

#![allow(unused)]
fn main() {
use goldy::types::{VertexBufferLayout, VertexFormat};

let layout = VertexBufferLayout::from_formats::<MyVertex>(&[
    VertexFormat::Float32x3, // position (12 bytes)
    VertexFormat::Float32x3, // normal   (12 bytes)
    VertexFormat::Float32x2, // uv       (8 bytes)
    VertexFormat::Uint32,    // color    (4 bytes)
]);
// stride = 36, 4 attributes
}

The builder panics if the summed format sizes don't equal size_of::<T>(), catching field-list mismatches at pipeline creation rather than producing silent GPU corruption.

Manual Layout

For full control, construct the layout directly:

#![allow(unused)]
fn main() {
use goldy::types::{VertexBufferLayout, VertexAttribute, VertexFormat};

let layout = VertexBufferLayout {
    stride: 32,
    attributes: vec![
        VertexAttribute { location: 0, format: VertexFormat::Float32x3, offset: 0 },
        VertexAttribute { location: 1, format: VertexFormat::Float32x3, offset: 12 },
        VertexAttribute { location: 2, format: VertexFormat::Float32x2, offset: 24 },
    ],
};
}

Empty Layout

When the vertex shader generates geometry from SV_VertexID (fullscreen triangles, instanced quads), use the default empty layout:

#![allow(unused)]
fn main() {
let pipeline = RenderPipeline::new(&device, &vs, &fs, &RenderPipelineDesc {
    vertex_layout: VertexBufferLayout::empty(),
    ..Default::default()
})?;
}

VertexBufferLayout::default() also returns an empty layout.

Vertex Formats

Available formats for vertex attributes:

FormatRust TypeSize
Float32f324
Float32x2[f32; 2]8
Float32x3[f32; 3]12
Float32x4[f32; 4]16
Uint32u324
Sint32i324
Uint8x4[u8; 4] (packed)4
Unorm8x4[u8; 4] (normalized)4

Vertex Data Flow

In Slang shaders, vertex attributes arrive through the [goldy_vertex] virtual entry point. The pipeline's VertexBufferLayout determines which attributes the hardware feeds into the shader's input struct. Attribute locations in the layout must match the shader's declared input locations.

For passes that bypass vertex buffers entirely, Slang helpers like vs_fullscreen_triangle() and quad_position() in goldy_exp.primitives generate geometry from SV_VertexID and SV_InstanceID.

Rendering Outputs

Windowed rendering uses present-on-scheme: SurfaceExchange + Transaction + Claim. Record copy or compute-to-present once in a retained Scheme; submit each frame and settle the claim.

All windowed Rust examples use this path.

SurfaceExchange

A SurfaceExchange wraps the platform window and records how scheme output reaches the swapchain:

#![allow(unused)]
fn main() {
use goldy::{SurfaceExchange, SurfaceConfig, PresentMode, DepthFormat};

let surface = SurfaceExchange::new(&ctx, &window)?;

// With explicit configuration and in-flight depth
let surface = SurfaceExchange::new_with_depth(
    &ctx,
    &window,
    3,
    SurfaceConfig {
        present_mode: PresentMode::Fifo,
        depth_format: Some(DepthFormat::Depth32Float),
    },
)?;
}

Depth testing uses an offscreen scheme-leased render target, not the swapchain drawable.

Bind helpers

MethodUse
bind_render_target(scheme, scene_rt)Offscreen render pass → surface copy
bind(scheme, texture)Texture → surface copy
bind_destination(scheme)Compute or other direct writes via with_present(&lease)

Each bind returns a reusable Transaction. After scheme.submit(), present with (&mut submission >> &transaction).take()?. The explicit &mut borrow is required by operator semantics and leaves other claims on the submission untouched. transaction.claim(&mut submission)? followed by claim.consume() remains for explicit multi-step settlement.

SurfaceConfig

#![allow(unused)]
fn main() {
pub struct SurfaceConfig {
    pub present_mode: PresentMode,
    pub depth_format: Option<DepthFormat>,
}
}
FieldPurposeDefault
present_modeVsync strategyAuto
depth_formatDepth buffer format, or None to disableNone

Present Modes

ModeBehaviorBackend Mapping
FifoVsync — wait for display refresh. No tearing, capped at monitor Hz.Metal displaySyncEnabled=YES, Vulkan FIFO, DX12 Present(1)
MailboxTriple-buffered — latest frame queued, older dropped. Low latency + no tearing.Vulkan MAILBOX. Falls back to Fifo on Metal and some DX12 configurations.
ImmediateNo sync, may tear. Maximum throughput for benchmarks.Metal displaySyncEnabled=NO, Vulkan IMMEDIATE, DX12 Present(0)
AutoGoldy chooses (Mailbox if available, then Fifo).—

Change the present mode at runtime:

#![allow(unused)]
fn main() {
surface.set_present_mode(PresentMode::Immediate)?;
let current = surface.present_mode();
}

Present-on-Scheme Frame Cycle

Record once at init (and on resize), submit each frame:

#![allow(unused)]
fn main() {
let mut pass = scheme.render_pass("main", &scene_rt, TargetLoad::Clear(Color::CORNFLOWER_BLUE));
pass.with_parcel(&vertex_buffer, NodeAccess::Read);
pass.set_pipeline(&pipeline);
pass.set_vertex_buffer(0, &vertices);
pass.draw(0..3, 0..1);
pass.finish();
let present = surface.bind_render_target(&mut scheme, &scene_rt)?;

// Each frame:
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
}

For pure compute-to-surface, use bind_destination and bind the returned lease in a compute node with with_present(&lease) instead of a render pass + copy.

Surface Queries

#![allow(unused)]
fn main() {
surface.width();
surface.height();
surface.size();        // (width, height)
surface.format();      // TextureFormat of the swapchain images
}

Always use surface.format() when creating pipelines to ensure a match:

#![allow(unused)]
fn main() {
let desc = RenderPipelineDesc {
    target_format: surface.format(),
    ..Default::default()
};
}

Resize Handling

Call resize() when the window size changes. Zero-size dimensions are silently ignored (common during window minimize). Rebuild the scheme when surface.size() changes.

SurfaceExchange::resize records the new extent immediately (and advances the pool generation) but defers the DXGI/ResizeBuffers work until the next drawable acquire. A burst of window-size events therefore only pays for one structural rebuild per presented frame.

#![allow(unused)]
fn main() {
surface.resize(width, height)?;
// rebuild scheme + transaction using surface.size()
}

Transaction Lifetime

  • Record a bind (bind_render_target, bind, or bind_destination) once when building the scheme.
  • Each frame: scheme.submit() then (&mut submission >> &transaction).take()?.
  • Each submission may be claimed at most once per transaction.
#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;
(&mut submission >> &present).take()?;
// claim consumed — do not reuse this submission's claim slot
}

Buffers

Buffer is a GPU memory allocation for storing typed data — uniforms, vertex data, index data, compute storage, or anything a shader needs to read or write.

For application-owned GPU memory, use Runtime and bind the returned Parcel in a scheme (with_parcel, set_vertex_buffer, MemoryExchange deposits). All Rust, Python, FFI, and .NET examples use this path.

#![allow(unused)]
fn main() {
use goldy::{BufferFlags, BufferKind};

let vertices = [/* Vertex2D ... */];
let vertex_parcel = runtime.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?;

// Uninitialized storage (e.g. a uniform updated each frame via MemoryExchange deposit):
let uniform = runtime.acquire_buffer_sized::<MyUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;
}

See runtime-owned-memory.md for textures, mosaics, and release.

With Raw Bytes

When the data is naturally &[u8], pass an explicit element stride to acquire_buffer:

#![allow(unused)]
fn main() {
use goldy::{BufferFlags, BufferKind};

// Stride defaults to 1 when omitted (byte-addressable)
let parcel = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    None,
    BufferFlags::empty(),
    Some(&raw_bytes),
)?;

// Explicit stride for structured buffer views
let parcel = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    Some(16),
    BufferFlags::empty(),
    Some(&raw_bytes),
)?;

// With flags (e.g. CPU_READABLE)
let parcel = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    Some(16),
    BufferFlags::CPU_READABLE,
    Some(&raw_bytes),
)?;
}

Empty Buffer

#![allow(unused)]
fn main() {
let parcel = runtime.acquire_buffer(
    4096,
    BufferKind::Scattered,
    None,
    BufferFlags::empty(),
    None,
)?;

// With a specific element stride
let parcel = runtime.acquire_buffer(
    4096,
    BufferKind::Scattered,
    Some(64),
    BufferFlags::empty(),
    None,
)?;
}

Low-level Runtime::alloc_* (crate-internal)

The runtime routes standalone allocations through VramAllocator via crate-internal Runtime::alloc_buffer helpers. Application code should not call these; use Runtime acquire APIs above.

Data Access Patterns

The access pattern describes how shader threads access the buffer. This drives hardware optimizations and determines the bindless descriptor category.

#![allow(unused)]
fn main() {
pub enum BufferKind {
    Scattered, // default — any thread, any address, read/write
    Broadcast, // all threads read the same address
}
}
PatternShader MappingUse When
ScatteredStructuredBuffer<T>, RWStructuredBuffer<T>General storage: particles, meshes, compute I/O
BroadcastConstantBuffer / uniform bufferUniform data: transforms, time, settings

For read-only input buffers that don't need write access, create with BufferKind::Scattered and access through goldy_buf_ro<T> in the shader. This enables hardware read-cache optimizations without requiring a separate access pattern.

BufferFlags

#![allow(unused)]
fn main() {
bitflags! {
    pub struct BufferFlags: u32 {
        const COPY_SRC      = 1 << 0;
        const COPY_DST      = 1 << 1;
        const CPU_READABLE  = 1 << 2;
        const CPU_WRITABLE  = 1 << 4;
    }
}
}
FlagPurpose
COPY_SRCBuffer can be a copy source
COPY_DSTBuffer can be a copy destination
CPU_READABLEPlacement hint: expect host claims on this parcel. Backends may keep the medium host-coherent so take() is wait + pointer. Semantics are identical without the flag (staged copy).
CPU_WRITABLEHost-mapped staging for deposits / upload copies. Prefer MemoryExchange::bind_deposit for application uploads.

Query RuntimeCapabilities::has_zero_copy_storage_readback to detect whether the backend honors CPU_READABLE with a mapped pointer.

Writing Data

Prefer MemoryExchange::bind_deposit for CPU→GPU uploads. Direct host writes on CPU_WRITABLE staging parcels remain for deposit/staging internals:

Raw bytes

#![allow(unused)]
fn main() {
buffer.write(offset, &bytes)?;
}

Typed data

#![allow(unused)]
fn main() {
buffer.write_data(offset, &[1.0f32, 2.0, 3.0])?;
}

Both methods write at a byte offset from the start of the buffer.

Reading Data

Use a host claim after submit:

#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;
let bytes = (&mut submission >> buffer.whole()).take::<u8>()?.to_vec();
}

Clearing

Zero-fill a region of the buffer:

#![allow(unused)]
fn main() {
buffer.clear(&device, offset, size)?;
}

Bindless Descriptors

Every buffer with Scattered or Broadcast access is registered in the global bindless descriptor set. Schemes bind parcels via with_parcel; the opaque ResourceHandle is available for identity / retention checks:

#![allow(unused)]
fn main() {
// Opaque typed identity — equality / hashing only; no public heap index
let handle = buffer.handle(ResourceAccess::Read).unwrap();

// Read-only SRV vs write UAV are distinct handles when both exist
let srv_handle = buffer.handle(ResourceAccess::Read).unwrap();
}

BufferView

A BufferView is a sub-region of an existing Buffer with its own bindless descriptor. The shader sees the sub-region as a zero-based buffer.

Creating Views

#![allow(unused)]
fn main() {
// Raw byte view — offset, size, optional element stride
let view = buffer.create_view(1024, 512, Some(16))?;

// Typed view — first element index, element count
let view = buffer.create_typed_view::<[f32; 4]>(0, 256)?;
}

Using Views

Views implement BufferSource, so they work anywhere a Buffer does — set_vertex_buffer, set_index_buffer, write_data, clear, and scheme parcel binding:

#![allow(unused)]
fn main() {
let view_handle = view.handle(ResourceAccess::Read).unwrap();
pass.set_vertex_buffer(0, &view);
}

Lifetime

Dropping a BufferView unregisters its descriptor but does not free the parent buffer's memory. Multiple views of the same buffer can exist simultaneously.

StructuredBufferElement

The StructuredBufferElement trait marks types safe for Runtime::acquire_buffer_with_data. It is implemented for common multi-byte primitives (u16, u32, f32, f64, etc.), fixed-size arrays of those types, and #[repr(C)] structs via #[derive(goldy_derive::StructuredBufferElement)].

Not implemented for u8/i8 — passing &[u8] would set stride to 1, which almost never matches the shader's expected struct stride. Use Runtime::acquire_buffer with an explicit element stride for raw bytes.

Rust-generated Slang structs

#[derive(goldy::GpuType)] makes Rust the source of truth for a structured-buffer element. Pass Type::GPU_TYPE when creating the shader and reference the type without redeclaring it in authored Slang. Declare logical fields only — Goldy packs to the Slang structured-buffer ABI at upload (acquire_buffer_with_data, write_data) and injects reserved __goldy_padN fields in generated Slang.

#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Copy, Clone, bytemuck::Pod, bytemuck::Zeroable, goldy::GpuType)]
struct Particle {
    position: [f32; 3],
    color: [f32; 4],
}

let shader = ShaderModule::from_slang_with_gpu_types(
    &device,
    source,
    &[Particle::GPU_TYPE],
)?;
let particles = device.acquire_buffer_with_data(&particles, BufferKind::Scattered)?;
}
// Particle is injected by Goldy.
[goldy_compute]
void cs_main(BufRO<Particle> particles, ThreadId id) {
    Particle particle = particles[id.x];
}

Do not bytemuck::bytes_of a GpuType into a GPU buffer: that is the host layout, not the storage ABI. Typed Goldy upload is the syscall.

The portable initial field set is f32, u32, i32, 2–4 lane arrays of those scalars, and square f32 matrices. Unsupported or sub-word host fields fail with an actionable error instead of silently changing the ABI.

Shared library modules can inject the same declarations via ShaderLibrary::from_source_with_gpu_types. Prefer that over redeclaring the struct in authored Slang. Stage I/O structs with semantics (POSITION, TEXCOORD*, SV_Position) stay authored in the shader.

Matrix Convention

Goldy uses column-major matrix layout in uniform/constant buffers across all backends. Rust math libraries (glam, nalgebra, ultraviolet) already store matrices column-major, so upload directly without transposing:

#![allow(unused)]
fn main() {
let uniforms = MyUniforms {
    projection: proj.to_cols_array_2d(),
    modelview: view.to_cols_array_2d(),
};
buffer.write_data(0, &[uniforms])?;
}

Goldy sets SLANG_MATRIX_LAYOUT_COLUMN_MAJOR at the Slang session level, so DX12, Vulkan, and Metal all interpret float4x4 the same way.

Runtime-Owned Memory

Runtime backs retained GPU memory — the same way a CPU program allocates from the process heap without asking for a separate heap handle. Acquire returns a Buffer (possibly partitioned) or a texture Parcel. Bind parcels, not raw aggregates — each parcel is one bindable unit (whole buffer, buffer range, or texture).

Context::release_buffer / Context::release_texture park a held parcel in that context's transient pool for epoch-gated reuse. There is no public RetainedPool type: acquisition is inherent on Runtime.

Quick start

#![allow(unused)]
fn main() {
use goldy::{BufferKind, BufferFlags, MemoryExchange, field, Init, NodeAccess, Scheme};

// Single-unit buffer (derefs to whole parcel):
let vertices = [/* ... */];
let vb = runtime.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?;

// Raw bytes with explicit stride:
let uniform_buf = runtime.acquire_buffer(
    raw_bytes.len() as u64,
    BufferKind::Scattered,
    Some(16),
    BufferFlags::empty(),
    Some(&raw_bytes),
)?;

// Uninitialized buffer (rewrite each frame with a MemoryExchange deposit):
let uniform = runtime.acquire_buffer_sized::<MyUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;

// Texture parcel:
let tex = runtime.acquire_texture(w, h, format, access, flags, Some(&pixels))?;

// Partitioned record (ping-pong, level geometry):
let cells = runtime.acquire_record([
    field("a", Init::data(&grid_a)),
    field("b", Init::zeros::<u32>(n)),
])?;
}

Scheme binding

#![allow(unused)]
fn main() {
let memory = MemoryExchange::new(&ctx);
let mut upload = Scheme::new(&ctx);
let deposit = memory.bind_deposit(&mut upload, goldy::DepositTarget::buffer_elements::<MyUniforms>(&*uniform, 1))?;
(&deposit << &data)?;
upload.submit()?;

let mut pass = scheme.render_pass("draw", &rt);
pass.with_parcel(&*vb, NodeAccess::Read);
pass.set_vertex_buffer(0, &*vb);
pass.draw(0..3, 0..1);

// Partitioned buffer: bind one field/range
pass.with_parcel(&cells["a"], NodeAccess::Read);

// Geometry bound via BufferSource only — register dependency without descriptor:
pass.with_buffer_dependency(&geometry, NodeAccess::Read);
}

Binding a multi-unit Buffer as one descriptor panics; index into fields instead.

Release

Call ctx.release_buffer(buffer) / ctx.release_texture(texture) when resizing or tearing down. While held, buffers need no epoch polling — the runtime stamps each parcel at submit.

Bindings

LanguageTypesAcquire
RustRuntime, Buffer, Parcelacquire_buffer*, acquire_record, acquire_texture
Pythongoldy.Runtime, goldy.Buffer, goldy.ParcelRuntime.acquire_buffer, Runtime.acquire_texture
C#Runtime, Buffer, Parcel, RecordBuilderRuntime.AcquireBuffer, Runtime.AcquireTexture
C / ffi-clientGoldyRuntime, GoldyBuffer, GoldyParcelgoldy_runtime_acquire_*

All examples under examples/ acquire from Runtime directly. See the bindings section for language-specific guides.

Textures and Samplers

Texture holds image data on the GPU. Sampler controls how that data is filtered and addressed when read in shaders. Together, they provide the standard texture sampling pipeline.

Creating a Texture

#![allow(unused)]
fn main() {
use goldy::{Texture, TextureKind, TextureFormat, TextureFlags};

let texture = Texture::new(
    &device,
    512, 512,
    TextureFormat::Rgba8Unorm,
    TextureKind::Interpolated,
    TextureFlags::COPY_DST,
)?;
}

With Initial Data

Data must be raw bytes matching width × height × bytes_per_pixel:

#![allow(unused)]
fn main() {
let pixels: Vec<u8> = load_image_rgba("sprite.png");
let texture = Texture::with_data(
    &device,
    &pixels,
    256, 256,
    TextureFormat::Rgba8Unorm,
    TextureKind::Interpolated,
    TextureFlags::COPY_DST,
)?;
}

Spatial Access Patterns

The access pattern determines how the texture is bound and accessed in shaders:

AccessShader MappingUse When
InterpolatedTexture2D with samplerImage data filtered between texels — sprites, materials, UI
DirectRWTexture2DStorage images, compute output, exact pixel reads/writes

Texture Formats

FormatBPPNotes
R8Unorm1Single-channel (masks, SDFs)
Rg8Unorm2Two-channel (normal maps, motion vectors)
Rgba8Unorm4Standard 8-bit RGBA
Rgba8UnormSrgb4sRGB color space
Bgra8UnormSrgb4sRGB, swapped channels (common swapchain format)
Bgra8Unorm4Linear, swapped channels
Rgba16Float8HDR
Rgba32Float16Full precision

TextureFlags

#![allow(unused)]
fn main() {
bitflags! {
    pub struct TextureFlags: u32 {
        const COPY_SRC       = 1 << 0;
        const COPY_DST       = 1 << 1;
        const RENDER_TARGET  = 1 << 2;
    }
}
}
FlagPurpose
COPY_SRCTexture can be a copy source (needed for host claims / GPU copies)
COPY_DSTTexture can be a copy destination (needed for deposits / copies)
RENDER_TARGETTexture can be used as a color attachment

Writing Data

Use MemoryExchange::bind_deposit with DepositTarget::texture for batched, non-blocking uploads:

#![allow(unused)]
fn main() {
let memory = MemoryExchange::new(&ctx);
let deposit = memory.bind_deposit(
    &mut scheme,
    DepositTarget::texture(&texture, 0, 0, width, height, pixels.len() as u64, 0),
)?;
(&deposit << pixels.as_slice())?;
}

For a one-shot fill at acquire time, pass init to [Runtime::acquire_texture].

Reading Data

Use a host claim after submit. The texture must have been created with TextureFlags::COPY_SRC:

#![allow(unused)]
fn main() {
let mut submission = scheme.submit()?;
let bytes = (&mut submission >> &texture).take::<u8>()?.to_vec();
}

Texture Queries

#![allow(unused)]
fn main() {
texture.width();
texture.height();
texture.format();
texture.byte_size();  // width * height * bytes_per_pixel
texture.access();     // TextureKind
texture.flags();      // TextureFlags
texture.is_owned();   // true if dropping destroys the GPU resource
}

Bindless Descriptors

Textures are registered in the global bindless descriptor set. The category depends on the access pattern: Interpolated maps to ResourceCategory::Texture, Direct maps to ResourceCategory::StorageImage. Schemes bind via with_parcel; the opaque handle is for identity checks:

#![allow(unused)]
fn main() {
let handle = texture.handle(ResourceAccess::Read).unwrap();
}

Texture Borrowing

Texture::borrow() creates a non-owning view that shares the GPU resource. Dropping a borrowed texture does not destroy the underlying resource. Use this when handing a texture reference into a system that may drop it before the owner is done.

#![allow(unused)]
fn main() {
let borrowed = texture.borrow();
assert!(!borrowed.is_owned());
// dropping `borrowed` does not free GPU memory
}

Depth Textures

Depth textures are created through SurfaceConfig or [Context::lease_render_target], not directly via Texture::new. Available depth formats:

FormatBitsStencil
Depth16Unorm16No
Depth24Plus24No
Depth24PlusStencil824 + 8Yes
Depth32Float32No
Depth32FloatStencil832 + 8Yes
#![allow(unused)]
fn main() {
let surface = SurfaceExchange::new_with_depth(
    &ctx,
    &window,
    3,
    SurfaceConfig {
        depth_format: Some(DepthFormat::Depth32Float),
        ..Default::default()
    },
)?;
}

Texture as Render Target

A texture created with TextureFlags::RENDER_TARGET can be used as a color attachment for offscreen rendering.

#![allow(unused)]
fn main() {
let offscreen = Texture::new(
    &device,
    1920, 1080,
    TextureFormat::Rgba16Float,
    TextureKind::Interpolated,
    TextureFlags::RENDER_TARGET | TextureFlags::COPY_SRC,
)?;
}

Samplers

A Sampler defines how texture coordinates are interpreted — filtering between texels and handling coordinates outside [0, 1].

Creating a Sampler

#![allow(unused)]
fn main() {
use goldy::{Sampler, SamplerDesc, FilterMode, AddressMode};

let sampler = Sampler::new(&device, &SamplerDesc {
    mag_filter: FilterMode::Linear,
    min_filter: FilterMode::Linear,
    mipmap_filter: FilterMode::Linear,
    address_mode_u: AddressMode::Repeat,
    address_mode_v: AddressMode::Repeat,
    ..Default::default()
})?;
}

Convenience Constructors

#![allow(unused)]
fn main() {
let nearest   = Sampler::nearest(&device)?;        // nearest filter, clamp to edge
let linear    = Sampler::linear(&device)?;          // linear filter, clamp to edge
let tiling    = Sampler::linear_repeat(&device)?;   // linear filter, repeat addressing
let default   = Sampler::default_sampler(&device)?; // nearest filter, clamp to edge
}

SamplerDesc

#![allow(unused)]
fn main() {
pub struct SamplerDesc {
    pub address_mode_u: AddressMode,    // default: ClampToEdge
    pub address_mode_v: AddressMode,    // default: ClampToEdge
    pub address_mode_w: AddressMode,    // default: ClampToEdge
    pub mag_filter: FilterMode,         // default: Nearest
    pub min_filter: FilterMode,         // default: Nearest
    pub mipmap_filter: FilterMode,      // default: Nearest
    pub max_anisotropy: f32,            // default: 1.0 (disabled)
    pub compare: Option<CompareFunction>, // default: None
    pub lod_min_clamp: f32,             // default: 0.0
    pub lod_max_clamp: f32,             // default: 32.0
}
}

Filter Modes

ModeEffect
NearestPixelated — nearest texel, no interpolation
LinearSmooth — bilinear interpolation between neighbors

Address Modes

ModeEffect for UVs outside [0, 1]
ClampToEdgeStretches the border texel
RepeatTiles the texture
MirrorRepeatTiles with alternating mirror flips

Depth Comparison Samplers

For shadow mapping and depth-based effects, set the compare field:

#![allow(unused)]
fn main() {
let shadow_sampler = Sampler::new(&device, &SamplerDesc {
    compare: Some(CompareFunction::LessEqual),
    mag_filter: FilterMode::Linear,
    min_filter: FilterMode::Linear,
    ..Default::default()
})?;
}

Bindless Descriptors

Samplers are registered under ResourceCategory::Sampler:

#![allow(unused)]
fn main() {
let handle = sampler.handle(ResourceAccess::Read).unwrap();
}

Binding Textures and Samplers in Shaders

Pass texture and sampler parcels through [ShaderResourceSlot] bindings:

#![allow(unused)]
fn main() {
use goldy::ShaderResourceSlot;

pass.with_shader_resources(&[
    ShaderResourceSlot::Parcel {
        parcel: &texture_parcel,
        access: NodeAccess::Read,
    },
    ShaderResourceSlot::Sampler(&sampler),
]);
pass.set_pipeline(&pipeline);
}

In Slang:

import goldy_exp;

[goldy_fragment]
float4 fs_main(Interpolated<float4> tex, Filter smp, float2 uv : TEXCOORD) {
    return tex.Sample(smp, uv);
}

Pooling and Sub-Allocation

GPU resource allocation is expensive. Creating many small buffers or textures each frame produces allocation overhead, descriptor churn, and VRAM fragmentation. Goldy routes client allocation through two doors:

DoorPermanenceAcquire
RuntimeCross-submission identity (deeds)acquire_texture / acquire_buffer / acquire_record
Context transient poolOne-submission tenancy (leases)Context::lease_texture / lease_buffer / lease_render_target, or acquire_transient_texture / acquire_transient_buffer

The runtime owns reclaim: retained release transfers into the transient pool with a ready_after stamp; transient bins reissue only after GPU retirement.

Partitioned retained buffers (one backing, many bindable fields) use internal scattered suballocation — see Runtime::acquire_record. Do not construct bump arenas or whole-object texture free-lists in client code.

Transient Allocation

Rendering pipelines allocate many short-lived GPU buffers and textures each frame — scratch storage, per-pass intermediates, filter pyramids. After submission those resources are dead until the GPU finishes, at which point the memory can be recycled. Clients must not poll timeline clocks to decide when reuse is safe.

Client door: TransientPool

Goldy exposes one transient door per Context:

AcquireReturn
Context::acquire_transient_bufferContext::return_transient_buffer
Context::acquire_transient_textureContext::return_transient_texture

Context leases (Context::lease_buffer / lease_texture / lease_render_target) realize through the same pool. Relinquished retained parcels enter via StampedParcel / ready_after; the pool reissues only after every stamped epoch has retired.

#![allow(unused)]
fn main() {
let scratch = ctx.acquire_transient_buffer(
    size,
    BufferKind::Scattered,
    BufferFlags::GPU_ONLY,
    Some(stride),
)?;
// ... bind, submit ...
ctx.return_transient_buffer(scratch);
}

See also Pooling and Sub-Allocation.

What was removed

The former public TransientAllocator strategies (BumpReset, Heap) and the internal scattered bump arena (BufferPool) were deleted: they had no in-tree consumers once in-tree callers moved to Runtime retained acquire / TransientPool. Whole-object epoch-gated recycle bins are the supported transient path; any future suballocation belongs behind that door, not as a parallel public API.

VRAM Allocator

All GPU buffer and texture allocations route through the runtime's internal allocator (pools call Runtime::alloc_*). Clients obtain bytes via Runtime / TransientPool; the allocator itself is not a public customization point.

Allocation Policy (Tracking and Budget)

Install a BudgetPolicy to track live GPU bytes and optionally enforce a cap:

#![allow(unused)]
fn main() {
use goldy::BudgetPolicy;
use std::sync::Arc;

let policy = Arc::new(BudgetPolicy::with_budget(512 * 1024 * 1024)); // 512 MiB
device.ensure_allocation_policy(policy)?;

println!("GPU memory in use: {} bytes", device.tracked_vram_bytes());
}

Use BudgetPolicy::new() when you only need telemetry without a hard budget.

Relationship to pools

  • Runtime retained acquire / TransientPool — recycling policy (when deeds and leases may be reissued after GPU retirement).
  • Runtime allocator + BudgetPolicy — provenance and optional byte budget.

Both doors allocate through the runtime, so an installed BudgetPolicy covers retained and transient parcels automatically.

Backend Architecture

Goldy ships three GPU backends today, each implemented natively against the platform graphics API — no translation layers (like MoltenVK) are involved. Two additional backends are in active development; a Tenstorrent backend is planned.

BackendStatusAPI LevelPlatformsRust Crate
VulkanShipped1.4+Windows, Linuxash
DX12ShippedDirect3D 12Windowswindows + gpu-allocator
MetalShippedTier 2+macOSmetal
CUDAIn progressCUDA Driver APINVIDIA GPUscudarc
WebGPUIn progressWebGPU (via wgpu)Cross-platformwgpu
CPUIn progress (compute-only)Slang host-callable JITHost—
TenstorrentPlannedTT-Metalium / TT-MLIRTenstorrent accelerators—

Native Implementations

Each backend maps Goldy concepts directly to the most natural primitives of its target API:

┌─────────────────────────────────────────────────────────────┐
│                    Goldy Core API                           │
│                                                             │
│   Runtime, Buffer, Texture, Pipeline, Scheme, ...          │
└─────────────────────────────────────────────────────────────┘
        │                    │                    │
        ▼                    ▼                    ▼
┌───────────────┐    ┌───────────────┐    ┌───────────────┐
│ Vulkan 1.4+   │    │ Metal 2+      │    │ DX12          │
│               │    │               │    │               │
│ • ash crate   │    │ • metal-rs    │    │ • windows-rs  │
│ • Dynamic     │    │ • Argument    │    │ • Root        │
│   rendering   │    │   buffers     │    │   signatures  │
│ • Descriptor  │    │ • Native      │    │ • Descriptor  │
│   indexing    │    │   hazard      │    │   heaps       │
│ • Buffer      │    │   tracking    │    │               │
│   device addr │    │               │    │               │
└───────────────┘    └───────────────┘    └───────────────┘

Translation layers introduce overhead from API mismatches, incompatible synchronization models, and extra validation. Native backends can leverage each API's strengths directly — for example, Metal's built-in hazard tracking, or Vulkan's descriptor indexing for bindless rendering.

Backend Selection

Default Selection

Goldy selects the platform-preferred backend automatically:

PlatformDefault Backend
macOSMetal
WindowsDX12
LinuxVulkan

Runtime Override — GOLDY_BACKEND

Override the backend at runtime with the GOLDY_BACKEND environment variable:

GOLDY_BACKEND=vulkan cargo run --example triangle
GOLDY_BACKEND=dx12   cargo run --example triangle

Accepted values (case-insensitive):

ValueBackendStatus
vulkan, vkVulkanShipped
dx12, d3d12, directxDX12Shipped
metal, mtlMetalShipped
cudaCUDAIn progress
webgpu, wgpuWebGPUIn progress

An unrecognized value produces a clear error listing the valid options.

Programmatic Selection

Query the active backend at runtime:

#![allow(unused)]
fn main() {
let instance = Instance::new()?;
println!("Backend: {:?}", instance.backend_type());
// Prints: Backend: Dx12   (on Windows)
// Prints: Backend: Vulkan (on Linux)
// Prints: Backend: Metal  (on macOS)
}

Compile-Time Selection (Feature Flags)

You can also restrict which backends are compiled in via Cargo features. This excludes both the code and the dependencies of unselected backends:

cargo build --no-default-features --features vulkan

See Conditional Compilation for details on feature flags, dependency exclusion, and CI setup.

Adapter Enumeration

After creating an Instance, enumerate available GPU adapters to inspect what hardware is present:

#![allow(unused)]
fn main() {
let instance = Instance::new()?;
let adapters = instance.enumerate_adapters();

for adapter in &adapters {
    println!("{}: {} ({})", adapter.id(), adapter.name(), adapter.vendor());
    println!("  Type: {:?}", adapter.device_type());
}
}

DeviceType

Each adapter reports a DeviceType:

VariantMeaning
DiscreteGpuDedicated graphics card with its own VRAM
IntegratedGpuGPU integrated into the CPU (shared memory)
CpuSoftware renderer (e.g. WARP on DX12, lavapipe on Vulkan)
OtherUnknown or unrecognized device class

Creating a Runtime

Request a device with a preferred DeviceType. If no adapter matches, Goldy falls back to the first available adapter:

#![allow(unused)]
fn main() {
let device = instance
    .request_adapter(&RequestAdapterOptions {
        power_preference: PowerPreference::HighPerformance,
        ..Default::default()
    })?
    .request_runtime(&RuntimeDescriptor::default())?;

// Or target a specific adapter by ID after enumeration:
let adapters = instance.enumerate_adapters();
let device = adapters[0].request_runtime(&RuntimeDescriptor::default())?;
}

Backend Capabilities

Runtime Capabilities

Query format preferences and backend-specific capabilities after creating a device:

#![allow(unused)]
fn main() {
let caps = device.capabilities();

println!("Surface format:     {:?}", caps.preferred_surface_format);
println!("Render target fmt:  {:?}", caps.preferred_render_target_format);
println!("Zero-copy readback: {}", caps.has_zero_copy_storage_readback);
}
CapabilityVulkanDX12Metal
Zero-copy CPU storage readbackYesNo (requires GPU copy to readback heap)Yes
Preferred surface formatBgra8UnormSrgbBgra8UnormSrgbBgra8UnormSrgb

Vulkan Backend

The Vulkan backend requires Vulkan 1.4+ and uses:

  • Dynamic rendering (VK_KHR_dynamic_rendering) — no VkRenderPass or VkFramebuffer objects
  • Descriptor indexing — bindless resource access by index in shaders
  • Buffer device address — 64-bit GPU pointers for direct memory access in shaders

DX12 Backend

The DX12 backend uses the windows crate and provides:

  • Root signatures for resource binding
  • Descriptor heaps for efficient bindless resource management
  • Shader compilation via Slang to DXIL
  • WARP software rasterizer for headless/CI use (GOLDY_DX12_FORCE_WARP=1)
  • GPU-Based Validation for deep debugging (GOLDY_DX12_GBV=1)

Metal Backend

The Metal backend uses the metal crate (native Metal, not MoltenVK):

  • Argument buffers for bindless resource binding
  • Native hazard tracking — Metal tracks resource hazards automatically
  • Shader compilation via Slang to Metal Shading Language

CUDA Backend (in progress)

Compute-focused prototype targeting NVIDIA GPUs via the CUDA Driver API (CUDA 13.1+ required for device-updatable graph nodes). Slang compiles to PTX; dispatches use the CUDA launch model. The cuda feature does not imply graphics. Buffer schemes, uploads/readbacks, timelines, indirect dispatch, and 2D textures/samplers (CUDA arrays + texture/surface objects) work.

Windows presentation: when cuda, graphics, and dx12 are all enabled, each CUDA device opens a LUID-matched DX12 companion. Surface frames expose an Rgba8Unorm shared scratch texture from a depth-3 staging ring independent of the DXGI swapchain image. Typical schemes (including Ekrano) write a CUDA-owned staging texture then CopyTexture into that imported scratch — the same local-then-copy pattern as native DX12 — before present's same-format CopyResource onto the R8G8B8A8_UNORM DXGI swapchain. CUDA signals a ready fence; DX12 waits it, copies, then signals a recycle fence. CUDA waits recycle only when wrapping the ring, so compute N+1 does not serialize behind present-copy N. Adapter mismatch, WARP, and linked-node adapters fail at device creation. A first-slice raster path is also available under the same feature gate: offscreen Rgba32Float and Rgba8Unorm render targets, indexed and non-indexed point/line/triangle pipelines (Slang → DXIL), bindless render bindings, optional DX12-only depth attachments / depth-stencil PSOs / ClearDepth, and CopyRenderTarget into present scratch / CUDA textures. Depth is not CUDA-imported (compute cannot sample it yet); stencil ops remain off. Vulkan interop is not supported.

Enable with the cuda Cargo feature (--no-default-features --features cuda auto-selects CUDA; in default builds use GOLDY_BACKEND=cuda):

cargo test --no-default-features --features cuda --test scheme_compute_integration
# Windows presentation:
GOLDY_BACKEND=cuda cargo run --example compute_to_surface --features examples

Texture notes for CUDA:

  • Sampled formats: R8Unorm, Rg8Unorm, Rgba8Unorm, Rgba8UnormSrgb, Rgba16Float, Rgba32Float. BGRA is rejected (no matching CUDA array swizzle).
  • Writable shader access (DirectSpatial<T>) supports size-matched pairs and Goldy’s typed-UAV emulation for convertible pairs:
    • DirectSpatial<float4> ↔ Rgba32Float (identity surface store)
    • DirectSpatial<float4> ↔ Rgba8Unorm (lazy PTX specialization: pack/unpack view over uint8_t4, DX12-style round(saturate(x)*255) on store). Partitions that launch this specialized variant stay on stream-replay segments between CUDA graph islands (or use full op-list retention when no graph-safe island remains).
    • DirectSpatial<half4> ↔ Rgba16Float
    • DirectSpatial<uint8_t4> ↔ Rgba8Unorm (Slang has no uchar4 alias)
    • Upload/copy/readback of other supported sampled formats still works.
  • Surfaces expose Rgba8Unorm imported scratch from a depth-3 ring (CUDA+DX12 interop tradeoff: extra staging textures so compute does not reuse frame N's scratch in N+1). Prefer writing a CUDA-owned Rgba8Unorm texture (or render target) and exporting with CopyTexture; direct launches onto imported scratch remain supported but are costlier under WDDM. The DXGI swapchain is matching R8G8B8A8_UNORM so present is a single CopyResource.
  • RuntimeCapabilities on CUDA advertise preferred_surface_format = Rgba8Unorm and preferred_render_target_format = Rgba8Unorm (no BGRA in supported lists).
  • CUDA has no separate sampler object — filtering is baked into each CUtexObject. A dispatch may use at most one distinct Filter configuration; additional distinct samplers are rejected.

Retainable partitions are split into alternating CUDA graph islands (contiguous graph-safe kernel launches) and stream-replayed boundary segments (clears, copies, format-specialized launches, present exports). Islands are captured on first submit and relaunched on clean resubmits; stream segments re-execute between them on the same CUDA stream. Indirect dispatches in graph islands use CUDA 13.1 device-updatable kernel nodes: an in-graph updater reads the GPU-resident DispatchShape and updates the consumer node's grid (or disables it for a zero / oversized shape). Uploads and other fully graph-unsafe partitions stay on the stream command-replay path, where indirect grids are resolved with a worker-side DtoH before cuLaunchKernel. Dynamic waits and completion events remain outside the captured graph. Stream capture is skipped when CUDA_LAUNCH_BLOCKING is set (including under GOLDY_VALIDATION=api).

With GOLDY_VALIDATION=api (or all), the CUDA backend enables Driver diagnostics: PTX JIT error/info logs on module load, host-side launch-limit checks, StructuredBuffer ABI checks, and per-op stream synchronize with labeled errors. It may set CUDA_LAUNCH_BLOCKING=1 when unset. Deep memory/race checking still requires external compute-sanitizer, not GOLDY_VALIDATION.

WebGPU Backend (in progress)

Cross-platform prototype built on wgpu. Intended for broader portability and browser-adjacent targets. Enable with the webgpu Cargo feature. Not yet at parity with the shipped Vulkan/DX12/Metal backends.

Compute buffers, scalar uniforms, indirect dispatch, and 2D textures/samplers work. Submit is non-blocking: the context timeline advances from wgpu's on_submitted_work_done callback (pumped by Runtime::poll). Host waits (Context::wait_until, host claims) block on the submission index, not on submit itself. Resources bind as a single @group(0) in shader-parameter order (no bindless heap). Texture notes:

  • Sampled formats: R8Unorm, Rg8Unorm, Rgba8Unorm, Rgba8UnormSrgb, Bgra8Unorm, Bgra8UnormSrgb, Rgba16Float, Rgba32Float (subject to adapter format features).
  • DirectSpatial<T> storage textures follow WGSL: the shader type encodes the format. Identity DirectSpatial<float4> is rgba32float. Goldy specializes packed 8-bit surfaces at dispatch: Rgba8Unorm → rgba8unorm, Bgra8Unorm → bgra8unorm (the latter needs wgpu BGRA8UNORM_STORAGE). sRGB formats are rejected for storage.
  • Uploads use queue.write_texture. Texture host-claim staging uses WebGPU's 256-byte row pitch; query_texture_copy_footprint reports the padded layout.
  • Surfaces: begin_frame acquires the wgpu drawable. Present picks Copy (same-format storage scratch → swapchain) when the scratch format can be a UAV (Rgba8Unorm, or Bgra8Unorm with BGRA8UNORM_STORAGE), otherwise Blit (Rgba8Unorm scratch + fullscreen pass for BGRA/sRGB). Direct compute-to-swapchain is not selected: wgpu 28 swapchain images only expose RENDER_ATTACHMENT, so storage bind groups fail even when the surface advertises STORAGE_BINDING. Override with GOLDY_WEBGPU_PRESENT=copy|blit. DirectSpatial<float4> shaders do not change; packed storage is specialized to the compute format (surface_format()). finish_present drops the acquired image and publishes SwapchainReturned. surface_resize reconfigures the swapchain and recreates scratch.
  • Raster: offscreen render targets, graphics PSOs (Slang → WGSL vs_main/fs_main), indexed and non-indexed draws, optional depth-stencil, and CopyRenderTarget. Vertex/fragment resources use the same packed @group(0) lowering as compute (no bindless heap). The webgpu feature implies graphics.
  • Compute sampling must use SampleLevel (WGSL has no implicit derivatives in compute).
cargo test --no-default-features --features webgpu --lib backend::webgpu

Tenstorrent Backend (planned)

Torus is a planned Fondaco runtime for Tenstorrent Tensix hardware. No implementation ships in Goldy today.

The GpuBackend Trait

All backends implement the GpuBackend trait, which defines the full interface for device management, resource creation, shader compilation, pipeline management, rendering, and compute dispatch:

#![allow(unused)]
fn main() {
pub trait GpuBackend: Send + Sync {
    fn backend_type(&self) -> BackendType;
    fn enumerate_adapters(&self) -> Vec<AdapterInfo>;
    fn create_device(&mut self, adapter_id: u32) -> Result<DeviceHandle>;
    fn create_buffer(&mut self, device: DeviceHandle, ...) -> Result<BufferHandle>;
    fn create_shader_with_paths(&mut self, device: DeviceHandle, ...) -> Result<ShaderHandle>;
    fn create_pipeline(&mut self, device: DeviceHandle, ...) -> Result<PipelineHandle>;
    // ... rendering, compute, surface, texture, sampler, timeline ...
}
}

Resources are identified by opaque u64 handles (DeviceHandle, BufferHandle, ShaderHandle, etc.) that each backend maps to native API objects internally.

Conditional Compilation

Most users should use GOLDY_BACKEND for runtime switching — see Backend Architecture.

Compile-time feature flags are useful when you need smaller binaries, faster builds, or want to verify that each backend compiles independently in CI.

When to Use Compile-Time Features

Use --no-default-features --features <backend> when you need:

  • Smaller binaries — exclude unused backend code
  • Faster builds — skip compiling heavy backend dependencies
  • Missing SDK — build on a system that lacks the Vulkan SDK or Windows SDK
  • CI matrix — verify each backend compiles independently
  • Compute-only builds — CUDA without raster, surfaces, or presentation

Feature Flags

Goldy defines one feature per backend, plus gpu, graphics, tensor, and instrumentation:

[features]
default = ["vulkan", "metal", "dx12", "instrumentation", "graphics", "tensor"]
graphics = ["dep:raw-window-handle"]
tensor  = []   # dense tensor algebra; does not imply graphics
gpu     = []   # implied by every real GPU backend (not mock)
vulkan  = ["dep:ash", "graphics", "gpu"]
dx12    = ["dep:windows", "dep:gpu-allocator", "dep:windows-core", "graphics", "gpu"]
metal   = ["dep:metal", "dep:cocoa", "dep:objc", "dep:core-graphics-types",
           "dep:foreign-types", "dep:block", "graphics", "gpu"]
cuda    = ["dep:cudarc", "gpu"]
webgpu  = ["dep:wgpu", "dep:pollster", "graphics", "gpu"]

instrumentation = ["dep:tracing-subscriber"]

graphics

graphics enables raster pipelines, render targets, surfaces, and presentation. Native backends (vulkan, dx12, metal) imply graphics, so enabling any of them keeps the full graphics+compute API.

Textures and samplers remain available without graphics — they are part of the GPGPU compute surface (storage images, sampling, copies, deposits, host claims).

gpu is an empty umbrella enabled by vulkan, dx12, metal, cuda, and webgpu. Use cfg(feature = "gpu") for tests that need Instance::new() rather than the always-compiled mock backend. Enabling gpu alone does not compile a backend.

cuda does not imply graphics.

Neither cude nor webgpu are a platform default (Metal / DX12 / Vulkan remain the defaults in normal builds). When you compile only cuda or webgpu — no native backend — Instance::new() selects that backend automatically. In a default multi-backend build, opt in with GOLDY_BACKEND=cuda or GOLDY_BACKEND=webgpu.

On Windows, enabling cuda together with graphics and dx12 (the usual case when adding cuda on top of default features) attaches a DX12 presentation companion to each CUDA device: LUID-matched DXGI adapter, shared float4 scratch textures, and swapchain present. The same gate enables a first-slice raster path (offscreen Rgba32Float targets, indexed/non-indexed point/line/triangle pipelines, bindless bindings, and optional DX12-only depth). Vulkan interop remains unsupported. Without that full gate, surface/present/raster APIs still return compute-only errors.

# CUDA compute-only
cargo test --no-default-features --features cuda --test scheme_compute_integration

# CUDA compute-only plus tensor algebra (no graphics)
cargo test --no-default-features --features cuda,tensor --test tensor

# CUDA + DX12 presentation + first-slice raster (Windows)
cargo check --no-default-features --features cuda,graphics,dx12
GOLDY_BACKEND=cuda cargo run --example compute_to_surface --features examples
cargo test --no-default-features --features cuda,graphics,dx12 --test cuda_dx12_raster
cargo test --no-default-features --features cuda,graphics,dx12 --test cuda_dx12_presentation
cargo test --no-default-features --features cuda,graphics,dx12 --test cuda_dx12_surface_lifecycle

Dependency Exclusion

Building with only one backend excludes both the code and the dependencies for the others:

FeatureDependencies
gpunone (umbrella; implied by each backend below)
vulkanash (+ graphics / raw-window-handle)
dx12windows, gpu-allocator, windows-core (+ graphics)
metalmetal, cocoa, objc, core-graphics-types, foreign-types, block (+ graphics)
cudacudarc
webgpuwgpu, pollster
# Default build on Windows — compiles Vulkan + DX12 dependencies
cargo build

# Vulkan-only build — downloads only ash (and enables graphics)
cargo build --no-default-features --features vulkan

# DX12-only build
cargo build --no-default-features --features dx12

# CUDA compute-only (no raster; surfaces need dx12+graphics on Windows)
cargo build --no-default-features --features cuda

# CUDA + DX12 presentation companion (Windows)
cargo build --no-default-features --features cuda,graphics,dx12

This can significantly reduce build times and binary size.

Platform-Specific Considerations

BackendAvailable OnNotes
vulkanWindows, Linux (any platform with a Vulkan loader)Broadest platform support; implies graphics
dx12Windows onlyGated by #[cfg(target_os = "windows")] — the feature is ignored on other platforms; implies graphics
metalmacOS onlyGated by #[cfg(target_os = "macos")] — the feature is ignored on other platforms; implies graphics
cudaAny platform with CUDA toolkitCompute prototype; on Windows with cuda+graphics+dx12, DX12 presentation companion + first-slice raster (Rgba32Float / Rgba8Unorm, indexed draws, DX12-only depth) are enabled. Does not imply graphics by itself. Vulkan interop still pending.
webgpuCross-platformvia wgpu; implies graphics

On macOS, the default backend is native Metal. Goldy does not require MoltenVK.

Default Features

The default feature set enables all three native backends plus instrumentation and graphics:

default = ["vulkan", "metal", "dx12", "instrumentation", "graphics", "tensor"]

To override, use --no-default-features and enable only what you need:

# Only Vulkan (graphics implied)
cargo build --no-default-features --features vulkan

# Vulkan + instrumentation
cargo build --no-default-features --features vulkan,instrumentation

# Metal-only on macOS
cargo build --no-default-features --features metal

# CUDA compute-only
cargo build --no-default-features --features cuda

FFI and Python Feature Passthrough

The goldy-ffi and goldy-py crates propagate features to the core goldy crate, so you can control backend selection in downstream builds. The same goldy-ffi build is consumed by C++, .NET, and goldy-ffi-client.

# FFI bindings with only Vulkan backend
cargo build -p goldy-ffi --no-default-features --features vulkan

# FFI with CUDA compute-only
cargo build -p goldy-ffi --no-default-features --features cuda

# Python bindings with only DX12 backend
cargo build -p goldy-py --no-default-features --features dx12

This is useful for creating platform-specific binary distributions.

Cross-Compilation

When cross-compiling, keep in mind that platform-gated features are silently ignored if the target platform doesn't match:

# Targeting macOS — dx12 feature is silently ignored, only metal + vulkan
# are active
cargo build --target aarch64-apple-darwin

# Targeting Windows — metal feature is silently ignored
cargo build --target x86_64-pc-windows-msvc --no-default-features --features dx12

For cross-compilation to work, you need the appropriate system SDKs available. Vulkan is the most portable backend since the ash crate only needs a Vulkan loader at runtime, not at compile time.

CI Matrix Example

Verify each backend compiles independently in CI:

# GitHub Actions
jobs:
  lint:
    strategy:
      matrix:
        include:
          - os: ubuntu-latest
            features: vulkan
          - os: windows-latest
            features: vulkan
          - os: windows-latest
            features: dx12
          - os: macos-latest
            features: metal
          - os: ubuntu-latest
            features: webgpu
          - os: windows-latest
            features: webgpu
          - os: macos-latest
            features: webgpu
    runs-on: ${{ matrix.os }}
    steps:
      - uses: actions/checkout@v4
      - run: cargo clippy --no-default-features --features ${{ matrix.features }} -- -D warnings

Checking the Active Backend

At runtime, query which backend was selected:

#![allow(unused)]
fn main() {
let instance = Instance::new()?;
println!("Backend: {:?}", instance.backend_type());
}

If no backend feature is enabled for the current platform, Instance::new() returns an error:

No GPU backend available — enable 'vulkan', 'dx12', 'metal', 'cuda', or 'webgpu'

In a default build (Vulkan + DX12 + Metal), use GOLDY_BACKEND=cuda or GOLDY_BACKEND=webgpu to opt into the in-progress compute prototypes.

Debugging and Observability

Goldy provides validation layers, structured instrumentation, and environment variable controls that together cover the full debugging workflow — from catching API misuse to profiling frame timing.

Validation

GOLDY_VALIDATION Environment Variable

The primary control for runtime validation. Accepts a comma-, semicolon-, or whitespace-separated list of categories:

ValueEffect
apiEnable backend GPU API validation (see below)
layoutEnable Rust ↔ Slang struct layout checks and buffer stride checks
host_accessPage-protect CPU-visible GPU copies (CPU backend parcels; stray host pointers fault)
scheme, graph, readbackHost-read staging checks, plus strict Accel build-before-trace in the same scheme
allEnable api, layout, timeline, scheme, and host_access
1, true, yesGPU API validation only (legacy shorthand; does not enable layout checks)

Categories can be combined:

# API validation only
GOLDY_VALIDATION=api cargo run --example triangle

# Layout validation only
GOLDY_VALIDATION=layout cargo run --example triangle

# Both
GOLDY_VALIDATION=all cargo run --example triangle
GOLDY_VALIDATION=layout,api cargo run --example triangle

API Validation

When GOLDY_VALIDATION includes api (or 1/true/yes), Goldy enables backend-specific validation:

BackendWhat Gets Enabled
VulkanVK_LAYER_KHRONOS_validation + VK_EXT_debug_utils at instance creation
MetalSets MTL_SHADER_VALIDATION=1 (if not already set) before the first device is created
DX12See DX12 Debug Layer below
WebGPUwgpu validation error scopes on shader/PSO create (always in debug builds; in release when GPU API validation is on) and on bind-group create (GPU API validation only)

On Vulkan, GOLDY_VALIDATION_FATAL=1 treats messenger ERROR records as hard failures (Err on later Goldy Result calls; panic on backend drop).

Goldy also validates the scheme graph on every Scheme::submit without an env var: dependency cycles, dispatch_mesh without set_mesh_pipeline, draw after a mesh pipeline, BLAS bound as a shader Accel, and BLAS geometry missing BufferFlags::ACCEL_INPUT. Failures are GoldyError::Validation strings that include a hint: with the fix. GOLDY_VALIDATION=scheme (or graph) additionally requires each Accel read to have a build_blas / build_tlas in the same scheme — turn that off if you built the AS in an earlier submit.

For Vulkan, validation is also enabled when VK_INSTANCE_LAYERS contains VK_LAYER_KHRONOS_validation (the standard loader-driven workflow).

Layout Validation

Layout validation catches mismatches between Rust struct layouts and their Slang shader counterparts at shader compile time, and buffer element-stride mismatches at dispatch time.

Enable via either:

GOLDY_VALIDATION=layout  cargo run
GOLDY_VALIDATE_LAYOUTS=1 cargo run   # legacy variable, equivalent

#[derive(LayoutCheckable)]

Annotate Rust structs that mirror Slang types to opt into automatic validation:

#![allow(unused)]
fn main() {
#[derive(LayoutCheckable)]
#[repr(C)]
struct SceneUniforms {
    projection: [[f32; 4]; 4],
    view: [[f32; 4]; 4],
    time: f32,
}
}

The derive macro generates a LAYOUT_CHECK constant containing the struct's name, total size, and per-field offsets. Pass it when creating a shader module:

#![allow(unused)]
fn main() {
let shader = ShaderModule::from_slang_with_options(
    &device,
    source,
    &[],          // extra search paths
    &[],          // defines
    Default::default(),
    &[SceneUniforms::LAYOUT_CHECK],
)?;
}

When layout validation is enabled, Goldy compiles the Slang shader, reflects each named struct, and compares:

  • Total struct size — Rust size_of vs. Slang reflection
  • Field offsets — each named field's byte offset

A mismatch produces an error naming the struct, the field, and the expected vs. actual offset — immediately surfacing padding or alignment bugs. When validation is disabled, the checks are skipped at zero cost.

Buffer Stride Validation

At dispatch time (when layout validation is enabled), Goldy also checks that each bound buffer's element_stride matches the stride the shader expects from Slang reflection. A mismatch produces an error like:

buffer element-stride mismatch in shader `my_shader`:
  slot 0: shader expects element stride 16 but buffer has 4

GOLDY_SHADER_VALIDATION — Static Shader Checks

Static checks over Slang's front-end IR, run at shader compile time. They have their own variable because they are a different kind of thing from the runtime categories above: each shader is compiled a second time to IR and analyzed whole-program (cost grows with shader size, not with what the app does), and a finding means "not proven", not "definitely wrong". For that reason GOLDY_VALIDATION=all does not turn them on.

The value is a list of check names processed left to right:

ValueEffect
boundsStatic bounds analysis: every dynamic array / vector / matrix index that cannot be proven inside 0 <= index < length is logged as a warning with its Slang source location and call path (interprocedural, across imported modules)
all (or 1/true/yes)Every check
-boundsRemove a check enabled earlier in the list (all,-bounds)
noneNothing
GOLDY_SHADER_VALIDATION=bounds cargo run --example triangle
GOLDY_VALIDATION=all GOLDY_SHADER_VALIDATION=all cargo test

Findings never fail a compile; they are logged once per distinct shader compile (including disk-cache hits) as shader validation (bounds): ....

DX12-Specific Debugging

DX12 Debug Layer

VariableValuesEffect
GOLDY_DX12_DEBUG1Force-enable the D3D12 debug layer (even in release builds)
GOLDY_DX12_NO_DEBUG1Disable the D3D12 debug layer (useful for parallel tests that crash the debug layer)
GOLDY_DX12_GBV1Enable GPU-Based Validation (very slow; requires the debug layer)

GPU-Based Validation (GBV) instruments shaders on the GPU to detect issues that the CPU-side debug layer cannot catch — such as out-of-bounds descriptor accesses and uninitialized resource reads. Expect a significant performance hit.

WARP Software Rasterizer

WARP is Microsoft's software implementation of D3D12. It runs on the CPU, so it works on headless CI runners with no GPU.

GOLDY_DX12_FORCE_WARP=1 cargo nextest run

After the first WARP device is created, Goldy prints a confirmation:

[WARP] d3d10warp.dll loaded from: C:\WINDOWS\SYSTEM32\d3d10warp.dll

On Windows, DX12 is the default backend, so GOLDY_DX12_FORCE_WARP=1 is the only variable you need to run tests on a machine without a GPU.

Structured Instrumentation

Goldy includes a structured instrumentation system built on the tracing crate. It provides named observation points with hierarchical dot-notation names and structured context data.

Enabling Instrumentation

Instrumentation requires the instrumentation Cargo feature (enabled by default). When disabled, all macros compile to no-ops at zero cost.

# Explicitly enable
cargo build --features instrumentation

# Disable (zero-cost removal)
cargo build --no-default-features --features vulkan

goldy_span! — Timed Sections

Create a span to measure the duration of a code section:

#![allow(unused)]
fn main() {
use goldy::goldy_span;

fn compile_shader(&self) {
    let _span = goldy_span!("slang.compile", target = "metal").entered();
    // ... compilation code ...
    // Duration is recorded automatically when _span is dropped
}
}

goldy_event! — Instant Markers

Emit a one-shot structured event:

#![allow(unused)]
fn main() {
use goldy::goldy_event;

goldy_event!("slang.library.load",
    path = %lib_path.display(),
    success = true
);
}

Built-in Observation Points

Goldy instruments its own internals at these observation points:

CategoryPoint NameEmitted Data
Slangslang.library.loadpath, success
slang.compile.starttarget, entry_points, bindless
slang.compile.endduration_ms, output_size, success
slang.reflection.extractparameter_blocks, fields
Shadershader.module.createbackend, shader_type
shader.pipeline.createpipeline_type, bind_groups
Resourceresource.buffer.createsize, usage
resource.texture.createdimensions, format
resource.bind_group.createbindings_count
Renderrender.frame.startframe_id
render.compute.dispatchworkgroups, pipeline
render.drawvertices, instances
render.frame.endframe_id, duration_ms

JSON Logging

Install a JSON file logger to capture all instrumentation output as structured JSON:

#![allow(unused)]
fn main() {
use goldy::instrumentation::install_json_logger;

install_json_logger("/tmp/goldy-debug.json")?;

// All subsequent goldy_span!/goldy_event! calls are written to the file
}

Filtering with RUST_LOG

Use the standard RUST_LOG environment variable to control verbosity. All Goldy instrumentation uses the goldy target:

RUST_LOG=goldy=debug cargo run --example triangle
RUST_LOG=goldy::render=trace cargo run --example triangle

Environment Variables Summary

VariableValuesEffect
GOLDY_BACKENDvulkan/vk, dx12/d3d12/directx, metal/mtlOverride backend selection
GOLDY_VALIDATIONapi, layout, host_access, all, 1/true/yesEnable validation categories
GOLDY_VALIDATION_FATAL1, true, yesFail on Vulkan Khronos ERROR messages
GOLDY_VALIDATE_LAYOUTS1, true, yesEnable layout validation (legacy; prefer GOLDY_VALIDATION=layout)
GOLDY_SHADER_VALIDATIONbounds, all, -bounds, noneStatic shader checks over Slang IR (separate from GOLDY_VALIDATION)
GOLDY_DX12_FORCE_WARP1Use WARP software rasterizer
GOLDY_DX12_DEBUG1Force-enable D3D12 debug layer in release
GOLDY_DX12_NO_DEBUG1Disable D3D12 debug layer
GOLDY_DX12_GBV1Enable GPU-Based Validation
GOLDY_SHADER_TIMING1Print Slang/PSO startup timings to stderr
GOLDY_CPU_SHADERS1Documented gate for debug CPU host-callable kernels (goldy::cpu_shaders)
RUST_LOGe.g. goldy=debugFilter instrumentation output

Common Debugging Patterns

Catch API misuse early

GOLDY_VALIDATION=api cargo run --example my_app

Turn on API validation during development to catch invalid GPU API calls. On Vulkan this enables the Khronos validation layer; on Metal it enables shader validation.

Diagnose struct layout bugs

GOLDY_VALIDATION=layout cargo test

If a LayoutCheckable struct diverges from its Slang counterpart (due to padding, alignment, or a field being added on only one side), the error message names the exact struct and field.

Catch stray CPU pointers into GPU copies

GOLDY_VALIDATION=host_access cargo test --test cpu_backend
GOLDY_VALIDATION=all cargo run

When host_access is on, the CPU backend allocates each parcel in its own page-aligned mapping (plus a guard page) and leaves it inaccessible except during deposit, kernel dispatch, and host claims. A leftover host pointer then faults instead of silently reading GPU-owned bytes. This is a debug allocator: slower, not complete (native device-local VRAM is not mapped), and meant to grow to staging buffers on other backends.

Headless CI on Windows

GOLDY_DX12_FORCE_WARP=1 cargo nextest run

WARP gives you a fully functional D3D12 device on machines with no GPU. Combine with GOLDY_VALIDATION=api for maximum coverage.

Profile frame timing

#![allow(unused)]
fn main() {
use goldy::instrumentation::install_json_logger;

install_json_logger("/tmp/goldy-profile.json")?;

// Run your application, then inspect the JSON output for
// render.frame.start / render.frame.end durations
}

Deep DX12 debugging

GOLDY_DX12_DEBUG=1 GOLDY_DX12_GBV=1 cargo run --example my_app

GPU-Based Validation catches GPU-side issues the CPU debug layer cannot see, at a significant performance cost. Use it when you suspect descriptor or resource access bugs.

CPU host-callable shaders (debug)

Goldy can JIT the same Slang compute kernels it runs on GPU and execute them on the host via Slang SLANG_SHADER_HOST_CALLABLE (getEntryPointHostCallable). This is an opt-in debug path so you can step a stage in a CPU debugger without maintaining a second handwritten Rust implementation.

Standalone compile/dispatch on host slices is the original debug path. Scheme submit on a compute-only CPU device is available with GOLDY_BACKEND=cpu. That backend is not a CPU renderer and not a replacement for Vulkan / DX12 / Metal / CUDA / lavapipe / WARP.

When to use it

  • Stepping a #[goldy::compute] / [goldy_compute] kernel in a native debugger
  • Checking buffer math on host slices before wiring GPU parcels
  • Replacing deleted CPU twins in clients (for example Ekrano) with the real Slang

How to run a kernel

#![allow(unused)]
fn main() {
use goldy::cpu_shaders::{self, CpuBinding};
use goldy::slang::SlangCompiler;

let compiler = SlangCompiler::new()?;
let kernel = cpu_shaders::compile_kernel(&compiler, &kernel_def, &["shaders"])?;
let mut data: Vec<u32> = (0..64).collect();
kernel.dispatch_1d(64, &mut [CpuBinding::u32s(&mut data)])?;
}

cpu_shaders::compile accepts [goldy_compute] source (or raw [shader("compute")] after you pack bindings yourself). The CPU wrapper keeps BufRO / Scattered as typed uniform entry-point parameters instead of Goldy bindless slot indices.

Set GOLDY_CPU_SHADERS=1 when you want the documented env gate (reserved for a future Runtime debug option). The compile APIs above are already opt-in; GPU paths ignore the variable.

Host-callable JIT uses vendored slang-llvm next to libslang. No extra C++ toolchain is required when that library is present. Do not set Slang SLANG_TARGET_FLAG_GENERATE_WHOLE_PROGRAM with the current vendored Slang: getEntryPointHostCallable SIGSEGVs. Goldy omits that flag.

What lowers

TypeCPU ABI
BufRO<T>, Scattered<T> (T = uint / int / float / bool){ T* data; size_t count }
Scalar uint / int / float / bool4-byte word
ThreadId, GroupThreadId, GroupIdSV_DispatchThreadID / SV_GroupThreadID / SV_GroupID
goldy_buf_len(buf)GetDimensions on the CPU structured buffer

Workgroups run serially through the Slang CPU prelude (ComputeVaryingInput start/end group IDs).

What does not lower yet

TypeNotes
Broadcast / gpu::Uniform<T> / constant-buffer structsNeeds a CPU constant-buffer view
ByteAddressCPU prelude has byte-address types; Goldy ABI packing is not wired
Interpolated<T> (sampled textures)No software texture path
DirectSpatial<T> (storage images)No software texture path
Filter / samplersTexture-only
[goldy_vertex] / [goldy_fragment]Compute only
Goldy bindless frame table (native wrapper)CPU uses the CUDA-shaped typed uniform preamble; scheme submit maps bindless indices onto host {data, count} views
Broadcast / textures / graphicsStill unsupported on GOLDY_BACKEND=cpu

Fine rasterization stays GPU-only until textures work. Interlocked / groupshared behavior follows the Slang CPU prelude (typically mutex or sequential atomics) and is not a substitute for GPU memory-model testing.

Python Bindings

Goldy provides Python bindings via PyO3, offering a Pythonic API for GPU programming with seamless NumPy integration.

Installation

From PyPI

pip install goldy

From Source

git clone https://github.com/koubaa/goldy.git
cd goldy/python
python -m venv .venv
source .venv/Scripts/activate   # platform-specific
pip install -e ".[dev]"

Slang is embedded when the extension is compiled; you do not run build-slang.py for local development. Rebuild after editing python/src/*.rs with maturin develop.

Requirements

  • Python 3.9+
  • NumPy 1.20+
  • A GPU with Vulkan 1.4+, DX12, or Metal Tier 2+ support (CUDA and WebGPU backends are in progress; Tenstorrent is planned)

Optional Dependencies

pip install goldy[dev]   # pytest, pillow
pip install pillow       # image output only

Quick Start

import goldy
import numpy as np

instance = goldy.Instance()
device = instance.request_adapter().request_runtime()
ctx = device.create_context()

vertices = np.array([
    0.0, -0.5, 1.0, 0.0, 0.0, 1.0,
    -0.5,  0.5, 0.0, 1.0, 0.0, 1.0,
     0.5,  0.5, 0.0, 0.0, 1.0, 1.0,
], dtype=np.float32)
vertex_parcel = device.acquire_buffer(vertices, goldy.BufferKind.SCATTERED)[0]

shader = goldy.ShaderModule.from_slang(device, goldy.Builtins.VERTEX_COLOR_2D)
pipeline = goldy.RenderPipeline(device, shader, shader, goldy.RenderPipelineDesc())

readback = device.acquire_texture(
    100, 100, goldy.TextureFormat.RGBA8_UNORM,
    goldy.TextureKind.DIRECT, copy_src=True, copy_dst=True,
)

scheme = goldy.Scheme(ctx)
rt = ctx.lease_render_target(100, 100, goldy.TextureFormat.RGBA8_UNORM)
with scheme.render_pass("triangle", rt, goldy.TargetLoad.clear(goldy.Color(0.1, 0.1, 0.2, 1.0))) as rp:
    rp.with_parcel(vertex_parcel, goldy.NodeAccess.READ)
    rp.set_pipeline(pipeline)
    rp.set_vertex_buffer_parcel(0, vertex_parcel)
    rp.draw(vertex_count=3)

scheme.copy_to_texture(rt, readback)
memory = goldy.MemoryExchange(ctx)
submission = scheme.submit()
pixels = np.frombuffer(submission.take_texture(readback), dtype=np.uint8).reshape(100, 100, 4)

NumPy Integration

Creating GPU Parcels from Arrays

vertices = np.array([
    # x, y, r, g, b, a
    0.0, -0.5, 1.0, 0.0, 0.0, 1.0,
    0.5,  0.5, 0.0, 1.0, 0.0, 1.0,
   -0.5,  0.5, 0.0, 0.0, 1.0, 1.0,
], dtype=np.float32)

parcel = device.acquire_buffer(vertices, goldy.BufferKind.SCATTERED)

Supported dtypes

NumPy dtypeTypical use case
np.float32Vertex positions, colors, uniforms
np.float64High-precision data
np.uint32Index buffers, compute data
np.int32Signed integer data
np.uint1616-bit index buffers
np.uint8Raw byte data

Reading Results Back to NumPy

Use SchemeSubmission.take / take_texture after submit:

submission = scheme.submit()
output = np.frombuffer(submission.take(parcel), dtype=np.float32)

Performance Tips

  • Create once, update often — avoid allocating new parcels every frame. Reuse retained buffers and update via upload schemes when needed.
  • Use np.float32 — match the GPU's expected dtype to avoid an extra conversion.
  • Ensure contiguity — sliced arrays may not be contiguous. Call np.ascontiguousarray() before uploading if needed.

Compute Shaders

Goldy supports GPU compute from Python using Slang shaders.

Basic Example

import goldy
import numpy as np

instance = goldy.Instance()
device = instance.request_adapter().request_runtime()
ctx = device.create_context()

data = np.arange(256, dtype=np.float32)
parcel = device.acquire_buffer(data, goldy.BufferKind.SCATTERED)[0]

SHADER = """
import goldy_exp;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<float> data, ThreadId id) {
    data[id.x] = data[id.x] * 2.0;
}
"""

shader = goldy.ShaderModule.from_slang(device, SHADER)
pipeline = goldy.ComputePipeline(device, shader)

scheme = goldy.Scheme(ctx)
scheme.node("double", pipeline).with_parcel(
    parcel, goldy.NodeAccess.READ_WRITE
).dispatch(4, 1, 1)
memory = goldy.MemoryExchange(ctx)
submission = scheme.submit()
output = np.frombuffer(submission.take(parcel), dtype=np.float32)

Ping-Pong Buffers

For iterative algorithms, alternate two buffer fields as input/output within one scheme (see python/examples/game_of_life.py).

Combining Compute and Graphics

Hybrid compute + render workflows use a single Scheme with both compute nodes and render passes (see python/examples/game_of_life.py and examples/game_of_life.rs).

Key Differences from Rust

AspectRustPython
Instance creationInstance::new()?goldy.Instance()
Error handlingResult<T, GoldyError>Raises goldy.GoldyError
Retained bufferruntime.acquire_buffer_with_data(&data, access)device.acquire_buffer(numpy_array, access) → Parcel
Render passscheme.render_pass(...)with scheme.render_pass(...) as rp:
Compute nodescheme.node(...).dispatch(...)scheme.node(...).with_parcel(...).dispatch(...)
Update retained buffer(&deposit << &data)?deposit << numpy_bytes
Resource lifetimeExplicit Arc<Runtime> ownershipManaged by Python GC via PyO3

Backend Selection

Goldy auto-selects the best backend per platform (DX12 on Windows, Vulkan on Linux). Override with GOLDY_BACKEND:

import os
os.environ["GOLDY_BACKEND"] = "vulkan"   # set before importing goldy

import goldy
instance = goldy.Instance()

API Reference

Core Classes

Instance

instance = goldy.Instance()
instance.backend_type            # BackendType (Vulkan, DX12, Metal; CUDA and WebGPU in progress)
instance.enumerate_adapters()    # list of AdapterInfo
instance.request_adapter()       # Adapter

Runtime acquire and Parcel

device = instance.request_adapter().request_runtime()
ctx = device.create_context()
parcel = device.acquire_buffer(data, access)  # data: numpy array or bytes
parcel.byte_size                            # int (bytes)

Tensor / TensorKernels

Dense tensor algebra records into the same Scheme. NumPy conversion copies through acquire and MemoryExchange — there is no zero-copy host view of GPU storage.

a = device.acquire_tensor(np.array([1.0, 2.0], dtype=np.float32), shape=[2])
b = device.zeros_tensor([2], goldy.TensorDType.F32)
kernels = goldy.TensorKernels(device)
scheme = goldy.Scheme(ctx)
kernels.fill_f32(scheme, "fill", b, 10.0)
c = kernels.add(scheme, "add", a, b)
memory = goldy.MemoryExchange(ctx)
submission = scheme.submit()
out = np.frombuffer(submission.take(c.parcel()), dtype=np.float32)

Scheme

scheme = goldy.Scheme(ctx)
rt = ctx.lease_render_target(w, h, goldy.TextureFormat.RGBA8_UNORM)

with scheme.render_pass("main", rt, goldy.TargetLoad.clear(goldy.Color.BLACK)) as rp:
    rp.with_parcel(buf, goldy.NodeAccess.READ)
    rp.set_pipeline(pipeline)
    rp.draw(vertex_count=3)

scheme.node("update", compute_pipeline).with_parcel(
    buf, goldy.NodeAccess.READ_WRITE
).dispatch(wg_x, wg_y, 1)

surface = goldy.SurfaceExchange.from_glfw(ctx, window)
present = surface.bind_render_target(scheme, rt)
submission = scheme.submit()
present.claim(submission).consume()

ShaderModule / RenderPipeline / ComputePipeline

Standard pipeline construction — see python/examples/triangle_headless.py.

Enums

goldy.DeviceType.DISCRETE_GPU | INTEGRATED_GPU | CPU | OTHER
goldy.TextureFormat.RGBA8_UNORM | RGBA8_UNORM_SRGB | BGRA8_UNORM
goldy.BufferKind.SCATTERED | BROADCAST
goldy.NodeAccess.READ | WRITE | READ_WRITE | OVERWRITE

Exceptions

All errors are raised as goldy.GoldyError:

try:
    device = instance.request_adapter().request_runtime()
except goldy.GoldyError as e:
    print(f"GPU error: {e}")

.NET Bindings

Goldy provides first-class C# bindings via P/Invoke interop over the native Rust FFI layer.

Installation

NuGet Package

dotnet add package Goldy

Or add to your .csproj directly:

<PackageReference Include="Goldy" Version="0.3.*" />

The NuGet package bundles native Goldy + Slang libraries for all supported platforms — no separate native installation is needed.

Building from Source

cargo build --package goldy-ffi --release
dotnet add reference path/to/goldy/dotnet/Goldy/Goldy.csproj

Requirements

  • .NET 8.0 or later
  • Windows x64, Linux x64, or macOS (x64 / arm64)
  • A GPU with Vulkan 1.4+, DX12, or Metal Tier 2+ support (CUDA and WebGPU backends are in progress; Tenstorrent is planned)

Quick Start

Headless Rendering

using Goldy;

using var instance = new Instance();
using var runtime = instance.RequestAdapter().RequestRuntime();
using var ctx = runtime.CreateContext();
using var readback = runtime.AcquireTexture(
    100, 100, TextureFormat.Rgba8Unorm, TextureKind.Direct,
    TextureFlags.CopySrc | TextureFlags.CopyDst);

using var scheme = new Scheme(ctx);
using var rt = scheme.LeaseRenderTarget(100, 100, TextureFormat.Rgba8Unorm);
using (var pass = scheme.RenderPass("clear", rt))
    pass.Clear(Color.CornflowerBlue);

scheme.CopyToTexture(rt, readback);
using var submission = scheme.Submit();
using var pixels = submission.Take(readback);

See Goldy.Examples/TriangleHeadless.cs for a full triangle readback demo.

Windowed Rendering

Record a retained scheme once, submit each frame, consume the present grant:

using var scheme = new Scheme(ctx);
var (sceneRt, present) = RecordScheme(scheme, swapchain, pipeline, vertexParcel, screen, bg);

using var submission = scheme.Submit();
present.Consume(submission);

See Goldy.Examples/TriangleWindow.cs and GameOfLifeWindow.cs.

Shaders (Slang)

Goldy uses Slang as its shader language across all backends:

var source = """
    [shader("vertex")]
    float4 vs_main(float2 pos : POSITION) : SV_Position {
        return float4(pos, 0.0, 1.0);
    }

    [shader("fragment")]
    float4 fs_main() : SV_Target {
        return float4(1.0, 0.5, 0.0, 1.0);
    }
    """;

using var shader = new ShaderModule(device, source);
using var pipeline = new RenderPipeline(device, shader, new RenderPipelineDesc
{
    TargetFormat = TextureFormat.Rgba8Unorm,
    Topology = PrimitiveTopology.TriangleList,
});

Resource Management

All Goldy objects implement IDisposable. Use using declarations or using blocks to ensure GPU resources are released promptly:

using var runtime = instance.RequestAdapter().RequestRuntime();
using var ctx = device.CreateContext();
using var scheme = new Scheme(ctx);

Key Differences from Rust

AspectRustC#
Instance creationInstance::new()?new Instance()
Error handlingResult<T, GoldyError>Exceptions
Runtime lifetimeArc<Runtime>IDisposable / using
Retained bufferruntime.acquire_buffer_with_data(&data, access)runtime.AcquireBuffer<T>(data, access) → Parcel
Submissionscheme.submit()?scheme.Submit() → SchemeSubmission
EnumsDeviceType::DiscreteGpuDeviceType.DiscreteGpu

API Reference

Scheme

public sealed class Scheme : IDisposable
{
    public Scheme(Context ctx);
    public SchemeComputeNodeScope ComputeNode(string label, ComputePipeline pipeline);
    public SchemeRenderTargetLease LeaseRenderTarget(uint width, uint height, TextureFormat format, ...);
    public SchemeRenderPassScope RenderPass(string label, SchemeRenderTargetLease lease);
    public void CopyToTexture(SchemeRenderTargetLease src, Texture dst);
    public void CopyToPresent(SchemeRenderTargetLease src, PresentLease dst);
    public SchemeSubmission Submit();
}

public sealed class MemoryExchange : IDisposable
{
    public MemoryExchange(Context ctx);
    public DepositTransaction BindDeposit(Scheme scheme, DepositTarget target);
}

public sealed class DepositTransaction : IDisposable
{
    public void Write(ReadOnlySpan<byte> data, ulong offset = 0);
    public static DepositTransaction operator <<(DepositTransaction deposit, byte[] data);
}

public sealed class SchemeSubmission : IDisposable
{
    public HostView Take(Parcel parcel);
    public HostView Take(Texture texture);
}

public sealed class HostView : IDisposable
{
    public int Length { get; }
    public ReadOnlySpan<byte> AsSpan();
    public byte[] ToArray();
}

SchemeRenderPassScope / SchemeComputeNodeScope

using (var pass = scheme.RenderPass("main", rt))
{
    pass.WithParcel(vertexParcel, NodeAccess.Read);
    pass.Clear(Color.CornflowerBlue);
    pass.SetPipeline(pipeline);
    pass.SetVertexBuffer(0, vertexParcel);
    pass.Draw(3);
}

using (var node = scheme.ComputeNode("update", computePipeline))
{
    node.WithParcel(stateBuf, NodeAccess.ReadWrite);
    node.Dispatch(wgX, wgY, 1);
}

SurfaceExchange / Transaction / Claim

using var surface = GlfwSurfaceExchange.Create(ctx, window);
var present = surface.BindRenderTarget(scheme, sceneRt);
// each frame:
using var submission = scheme.Submit();
present.Claim(submission).Consume();

Graphics and compute both go through Scheme.

Enums

public enum DeviceType   { DiscreteGpu, IntegratedGpu, Cpu, Other }
public enum BackendType  { Vulkan, Metal, Dx12 }  // CUDA and WebGPU in progress in core Goldy
public enum BufferKind   { Scattered, Broadcast }
public enum NodeAccess   { Read, Write, ReadWrite, Overwrite }

Headless vs windowed submission

Headless: record a scheme, Submit(), then submission.Take(parcel) / Take(texture).

Windowed: record once with SurfaceExchange.BindRenderTarget (or BindDestination for compute-to-surface); each frame call Submit(), then transaction.Claim(submission).Consume().

C++ Bindings

Goldy provides C and C++ bindings over the native goldy-ffi library. The C++ layer (goldy.hpp) wraps the auto-generated C API (goldy.h) with RAII types and exceptions.

Installation

vcpkg

# Add to your vcpkg.json
{
    "dependencies": ["goldy"]
}

# Or install directly
vcpkg install goldy

Conan

# conanfile.txt
[requires]
goldy/0.3.0

Building from Source

# Build the native library (requires Rust: https://rustup.rs)
cargo build --package goldy-ffi --release

# Configure and build examples
cd cpp
cmake -B build -DGOLDY_BUILD_FROM_SOURCE=ON
cmake --build build --target triangle_headless

On Windows, if MSVC cannot find standard headers, use x64 Native Tools Command Prompt for VS 2022 or run cpp/build.bat, which sets up the MSVC environment before invoking CMake.

Requirements

  • C++20 compiler (MSVC 2019+, GCC 10+, Clang 12+)
  • Rust toolchain (for building goldy-ffi from source)
  • A GPU with Vulkan 1.4+, DX12, or Metal Tier 2+ support (CUDA and WebGPU backends are in progress; Tenstorrent is planned)
  • Slang is embedded in the Goldy build — no separate SDK install for normal use

Native Library Deployment

The goldy_ffi shared library and Slang runtime DLLs must be on the loader path at runtime. CMake post-build steps copy them next to example binaries when building from source. For your own applications, ship goldy_ffi.dll / libgoldy_ffi.so / libgoldy_ffi.dylib alongside your executable, with the Slang shared libraries in the same directory.

Quick Start

Headless Rendering

#include <goldy.hpp>

#include <cstdint>
#include <iostream>

struct Vertex {
    float position[2];
    float color[4];
};

int main() {
    try {
        goldy::Instance instance;
        goldy::Runtime device = instance.request_adapter().request_runtime();
        goldy::Context ctx(device);

        const Vertex vertices[] = {
            {{0.0f, -0.5f}, {1.0f, 0.0f, 0.0f, 1.0f}},
            {{-0.5f, 0.5f}, {0.0f, 1.0f, 0.0f, 1.0f}},
            {{0.5f, 0.5f}, {0.0f, 0.0f, 1.0f, 1.0f}},
        };

        goldy::Buffer vertex_buffer = device.acquire_buffer_with_data(
            std::span<const Vertex>(vertices),
            goldy::BufferKind::Scattered);

        goldy::ShaderModule shader(device, goldy::ShaderModule::builtin_vertex_color_2d());

        GoldyVertexAttribute attributes[] = {
            {0, GOLDY_VERTEX_FORMAT_FLOAT32X2, 0},
            {1, GOLDY_VERTEX_FORMAT_FLOAT32X4, static_cast<uint32_t>(sizeof(float) * 2)},
        };

        GoldyRenderPipelineDesc desc{};
        desc.vertex_attributes = attributes;
        desc.vertex_attribute_count = static_cast<uint32_t>(std::size(attributes));
        desc.vertex_stride = sizeof(Vertex);
        desc.topology = GOLDY_PRIMITIVE_TOPOLOGY_TRIANGLE_LIST;
        desc.target_format = GOLDY_TEXTURE_FORMAT_RGBA8_UNORM;

        goldy::RenderPipeline pipeline(device, shader, shader, desc);

        GoldyTextureFlags readback_flags{};
        readback_flags._0 = goldy::TextureFlags::CopySrc | goldy::TextureFlags::CopyDst;
        goldy::Texture readback = device.acquire_texture(
            800, 600, GOLDY_TEXTURE_FORMAT_RGBA8_UNORM,
            GOLDY_TEXTURE_KIND_DIRECT, readback_flags);

        goldy::Scheme scheme(ctx);
        goldy::SchemeRenderTargetLease rt = ctx.lease_render_target(
            800, 600, GOLDY_TEXTURE_FORMAT_RGBA8_UNORM, nullptr);
        {
            auto pass = scheme.render_pass("triangle", rt, goldy::TargetLoad::clear(goldy::Color::cornflower_blue()));
            pass.with_field(vertex_buffer, 0, goldy::NodeAccess::Read)
                .set_pipeline(pipeline)
                .set_vertex_buffer(0, vertex_buffer)
                .draw(0, 3);
        }
        scheme.copy_to_texture(rt, readback);
        goldy::SchemeSubmission submission = scheme.submit();
        goldy::HostView view = submission.take(readback);
        std::cout << "Rendered " << view.size() << " bytes\n";
        return 0;
    } catch (const goldy::Exception& e) {
        std::cerr << "Goldy error: " << e.what() << '\n';
        return 1;
    }
}

See cpp/examples/triangle_headless.cpp for the full example.

Windowed Rendering

Use goldy::SurfaceExchange for swapchain presentation. See cpp/examples/triangle.cpp (Win32 / macOS).

Shaders (Slang)

Goldy uses Slang as its shader language across all backends:

const char* source = R"(
import goldy_exp;

[goldy_vertex]
float4 vs_main(Vertex2D v) : SV_Position {
    return float4(v.position, 0.0, 1.0);
}

[goldy_fragment]
float4 fs_main(Vertex2D v) : SV_Target {
    return float4(v.color);
}
)";

goldy::ShaderModule shader(device, source);

Resource Management

All C++ wrapper types use RAII — destructors release GPU handles automatically. Operations that can fail throw goldy::Exception:

try {
    goldy::Instance instance;
    // ...
} catch (const goldy::Exception& e) {
    std::cerr << "Goldy error: " << e.what() << "\n";
}

Key Differences from Rust

AspectRustC++
Instance creationInstance::new()?goldy::Instance instance
Error handlingResult<T, GoldyError>goldy::Exception
Runtime lifetimeRuntime (cheap Clone)RAII destructor
Retained bufferruntime.acquire_buffer_with_data(&data, access)runtime.acquire_buffer_with_data(span, access)
Render passscheme.render_pass(...)scheme.render_pass(...) (RAII scope)
Readbackclaim.consume(&submission)submission.take(parcel) / submission.take(texture)

API Reference

Core Classes

ClassDescription
goldy::InstanceEntry point, adapter enumeration
goldy::Runtime / goldy::ContextMachine root and submission timeline
goldy::RecordBuilderPartitioned buffer records (ping-pong fields)
goldy::SchemeRetained dependency graph
goldy::MemoryExchangeCPU→GPU deposit
goldy::SurfaceExchangeWindow swapchain (Win32 / macOS / Wayland)
goldy::ShaderModuleCompiled Slang shader
goldy::RenderPipeline / goldy::ComputePipelineGraphics/compute pipelines
goldy::SamplerTexture sampler

Scheme

goldy::Scheme scheme(ctx);
goldy::SchemeRenderTargetLease rt = ctx.lease_render_target(w, h, format, nullptr);

{
    auto pass = scheme.render_pass("main", rt, goldy::TargetLoad::clear(color));
    pass.with_field(buf, 0, goldy::NodeAccess::Read)
        .set_pipeline(pipeline)
        .set_vertex_buffer(0, buf)
        .draw(0, 3);
}

auto node = scheme.compute_node("update", compute_pipeline);
node.with_field(buf, 0, goldy::NodeAccess::ReadWrite)
    .dispatch(wg_x, wg_y, 1);

goldy::SchemeSubmission submission = scheme.submit();

MemoryExchange / SurfaceExchange

goldy::SchemeSubmission submission = scheme.submit();
goldy::HostView pixels = submission.take(texture);

goldy::SurfaceExchange surface(ctx, window_handle, width, height);
auto present = surface.bind_render_target(scheme, rt);
goldy::SchemeSubmission submission = scheme.submit();
present.claim(submission).consume();

goldy::MemoryExchange memory(ctx);
auto deposit = memory.bind_deposit(scheme, goldy::DepositTarget::buffer(parcel, capacity));
deposit << std::vector<uint8_t>{1, 2, 3, 4};
scheme.submit();

Raw C API

For C code or when you need low-level control, use goldy.h directly. Failed calls return null or error codes; call goldy_get_last_error() for details:

#include <goldy.h>

GoldyInstance* instance = goldy_instance_create();
if (!instance) {
    const char* error = goldy_get_last_error();
    // handle error
}

GoldyAdapterInfo info = {};
goldy_instance_get_adapter(instance, 0, &info);
GoldyRuntime* device = goldy_instance_create_runtime_for_adapter(instance, info.id);
// ...

goldy_runtime_destroy(device);
goldy_instance_destroy(instance);

The tensor feature (on by default) adds goldy_runtime_acquire_tensor, goldy_tensor_kernels_*, goldy_tensor_add, goldy_tensor_matmul, and goldy_tensor_fill_f32. C++ wraps them as goldy::Tensor and goldy::TensorKernels.

Platform Support

PlatformHeadless SchemeWindowed Surface
Windows x64YesYes
Linux x64YesYes (Wayland; X11 not supported)
macOS x64 / ARM64YesYes

Backend Selection

Goldy auto-selects the best backend per platform. Override with GOLDY_BACKEND (set before creating an Instance):

GOLDY_BACKEND=vulkan ./my_app

When building goldy-ffi for a specific platform, pass backend features through:

cargo build -p goldy-ffi --no-default-features --features vulkan

Examples

ExampleDescription
cpp/examples/triangle_headless.cppOffscreen triangle + readback
cpp/examples/triangle.cppWindowed triangle (GLFW)
cpp/examples/compute_simple.cppCompute dispatch
cpp/examples/game_of_life.cppHybrid compute + render

Rust FFI Client

goldy-ffi-client is a Rust crate that loads the goldy-ffi native library at runtime and exposes the same RAII API as the core goldy crate. Instead of statically linking the goldy library, it calls the stable C ABI through libloading (LoadLibrary on Windows, dlopen on Unix).

This is the same native boundary used by the C++ and .NET bindings. Python is different — it links the core goldy crate directly via PyO3.

When to Use

Use caseCrate
Normal Rust applicationsgoldy (static link, published on crates.io)
FFI integration testsgoldy-ffi-client
Validating the C ABI from Rustgoldy-ffi-client
Swapping the native library without recompiling the clientgoldy-ffi-client

The ffi-client API mirrors the core Rust crate: Instance, Runtime, Scheme, MemoryExchange, SurfaceExchange, Tensor / TensorKernels (behind the tensor feature), and the rest of the Fondaco programming model are available with the same names and patterns.

Installation

goldy-ffi-client is a workspace crate — it is not published to crates.io. Add it as a path dependency:

[dependencies]
goldy-ffi-client = { path = "../ffi-client" }

Build the native library first:

cargo build -p goldy-ffi

Then build or run ffi-client examples:

cd ffi-client
cargo run --example triangle_headless

Requirements

  • Rust 2021 edition
  • A built goldy_ffi shared library (goldy_ffi.dll / libgoldy_ffi.so / libgoldy_ffi.dylib)
  • A GPU with Vulkan 1.4+, DX12, or Metal Tier 2+ support (CUDA and WebGPU backends are in progress; Tenstorrent is planned)

Library Discovery

At runtime, ffi-client searches for the native library in this order:

  1. GOLDY_FFI_PATH — full path to the goldy_ffi dylib
  2. GOLDY_FFI_LIB_DIR — compile-time directory from the goldy-ffi build
  3. The directory containing the running executable

On Windows, ffi-client also calls SetDllDirectoryW so Slang DLLs next to goldy_ffi.dll are found.

# Point at a specific build of the native library
GOLDY_FFI_PATH=/path/to/libgoldy_ffi.so cargo run --example triangle_headless

Quick Start

Headless Rendering

use goldy_ffi_client::{
    shader::builtins, BufferKind, Color, Context, RuntimeDescriptor, Instance, NodeAccess,
    RenderPipeline, RenderPipelineDesc, RequestAdapterOptions, Scheme,
    ShaderModule, TargetLoad, TextureFlags, TextureFormat, TextureKind, Vertex2D,
};

fn main() -> goldy_ffi_client::Result<()> {
    let instance = Instance::new()?;
    let device = instance
        .request_adapter(&RequestAdapterOptions::default())?
        .request_runtime(&RuntimeDescriptor::default())?;
    let ctx = Context::new(&device)?;

    let vertices = [
        Vertex2D { position: [0.0, -0.5], color: [1.0, 0.0, 0.0, 1.0] },
        Vertex2D { position: [-0.5, 0.5], color: [0.0, 1.0, 0.0, 1.0] },
        Vertex2D { position: [0.5, 0.5], color: [0.0, 0.0, 1.0, 1.0] },
    ];
    let vertex_buffer = device.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?;

    let readback = device.acquire_texture(
        64, 64, TextureFormat::Rgba8Unorm, TextureKind::Direct,
        TextureFlags::COPY_SRC.union(TextureFlags::COPY_DST), None,
    )?;

    let shader = ShaderModule::from_slang(&device, builtins::VERTEX_COLOR_2D)?;
    let pipeline = RenderPipeline::new(
        &device, &shader, &shader,
        &RenderPipelineDesc {
            vertex_layout: Vertex2D::layout(),
            target_format: TextureFormat::Rgba8Unorm,
            ..Default::default()
        },
    )?;

    let mut scheme = Scheme::new(&ctx)?;
    let rt = ctx.lease_render_target(64, 64, TextureFormat::Rgba8Unorm, None)?;
    {
        let mut pass = scheme.render_pass("triangle", &rt, TargetLoad::Clear(Color::BLACK));
        pass.with_buffer(&vertex_buffer, NodeAccess::Read);
        pass.set_pipeline(&pipeline);
        pass.set_vertex_buffer(0, &vertex_buffer);
        pass.draw(0..3, 0..1);
        pass.finish_recorded();
    }
    scheme.copy_to_texture(&rt, &readback)?;

    let memory = goldy_ffi_client::MemoryExchange::new(&ctx)?;
    let mut submission = scheme.submit()?;
    let pixels = submission.take_texture(&readback)?;

    println!("Rendered {} bytes", pixels.len());
    Ok(())
}

See ffi-client/examples/triangle_headless.rs for the full example.

Windowed Rendering

See ffi-client/examples/triangle.rs and ffi-client/examples/game_of_life.rs (winit).

Compute

#![allow(unused)]
fn main() {
use goldy_ffi_client::{ComputePipeline, Context, Instance, MemoryExchange, NodeAccess, Scheme, ShaderModule};

let mut scheme = Scheme::new(&ctx)?;
let mut node = scheme.compute_node("double", &pipeline);
node.with_buffer(&buf, NodeAccess::ReadWrite);
node.dispatch(1, 1, 1);

let memory = MemoryExchange::new(&ctx)?;
let parcel = buf.field(0)?;
let deposit = memory.bind_deposit(&mut scheme, goldy_ffi_client::DepositTarget::buffer(&parcel, 16))?;
(&deposit << &[1u8, 2, 3, 4])?;
let mut submission = scheme.submit()?;
let bytes = submission.take(&parcel)?;
}

See ffi-client/examples/compute_simple.rs.

Resource Management

All ffi-client types use RAII via Drop. Errors are returned as goldy_ffi_client::Result<T> with GoldyError — the same pattern as the core goldy crate.

Key Differences from Core goldy

Aspectgoldygoldy-ffi-client
LinkingStatic (compiled into your binary)Dynamic (libloading at runtime)
Distributioncrates.ioWorkspace path dependency
API surfaceReference implementationMirrors core API over C ABI
Native libraryEmbedded in your binarySeparate goldy_ffi dylib required
Crate namegoldygoldy_ffi_client

Functionally, application code looks nearly identical. The main difference is build and deployment: ffi-client binaries need the goldy_ffi shared library (and Slang DLLs) available at runtime.

Backend Selection

GOLDY_BACKEND works the same as with the core crate — set it before creating an Instance:

GOLDY_BACKEND=vulkan cargo run --example triangle_headless

When building goldy-ffi, pass backend features through:

cargo build -p goldy-ffi --no-default-features --features vulkan

Examples

ExampleDescription
ffi-client/examples/triangle_headless.rsOffscreen triangle + readback
ffi-client/examples/triangle.rsWindowed triangle (winit)
ffi-client/examples/compute_simple.rsCompute dispatch
ffi-client/examples/game_of_life.rsHybrid compute + render (windowed)
ffi-client/examples/game_of_life_headless.rsGame of Life readback

Examples Gallery

Goldy ships 24 Rust examples, each a complete runnable program. Every example has a page here with a recording of it running, plus its Rust and Slang source inlined straight from the repository — so what you watch is what the code does, and what you read is what compiles.

The clips are not screen-captured from a desktop window. scripts/record_example_captures.sh runs each example headlessly: the example writes packed RGBA pixels, and ffmpeg stitches them into a WebM. Rerun it after changing an example's visuals.

Running Examples

cargo run --features examples --example <name> --release

All windowed examples exit on Escape, handle window resizes, and auto-exit after a soak period that GOLDY_EXAMPLE_TIMEOUT overrides. To run every example back to back:

./run_all_examples.sh
GOLDY_BACKEND=webgpu ./run_all_examples.sh
EXAMPLE_TIMEOUT=10 ./run_all_examples.sh

Backends

The examples run on every shipped backend — Vulkan 1.4+, DX12, and Metal Tier 2+ — and on the in-progress WebGPU backend, selected with GOLDY_BACKEND (see Backend Architecture):

GOLDY_BACKEND=webgpu cargo run --no-default-features --features webgpu,examples --example triangle

Two examples probe capabilities and exit cleanly when they are missing: mesh_triangle needs mesh shaders and ray_query needs ray query. Capture skips writing a clip when the adapter used for recording lacks the capability.

Headless capture does not present to a window, so it works on Vulkan (including lavapipe) without a Wayland or X11 display. GOLDY_BACKEND still selects the backend when you want WebGPU or CUDA instead.

Bindless Basics

Fundamental Goldy patterns: vertex buffers, surfaces, uniforms, and fragment shaders.

ExampleWhat it demonstrates
triangleMinimal windowed program: retained scheme, offscreen render pass, present via SurfaceExchange.
mesh_triangleThe triangle present path driven by MeshPipeline and dispatch_mesh.
gradientAnimated fullscreen gradient from a time uniform, with vertex-less rendering.
checkerboardProcedural animated checkerboard via UV distortion in a fragment shader.

Compute Workflows

ComputePipeline and Scheme for GPU-side data processing, including compute-to-surface.

ExampleWhat it demonstrates
compute_particlesCompute updates particle positions; graphics renders instanced quads.
game_of_lifeConway's Game of Life with ping-pong sub-views in one retained mosaic parcel.
compute_to_surfacePure compute rendering — no RenderPipeline, writes the drawable directly.
ray_queryTriangle BLAS/TLAS with inline RayQuery in [goldy_compute].
tensor_algebraHeadless dense tensors: fill, broadcast, and GEMV via TensorRecorder.

Graphics Pipelines

Classic rendering techniques: depth testing, textures, instancing, and 3D projection.

ExampleWhat it demonstrates
solid_cubeSolid 3D cube with per-face colours and a depth buffer.
spinning_cube3D wireframe cube using line primitives.
depth_quadsDepth buffer proves draw-order independence.
textured_quadProcedural texture on a quad with stage-local resources.
instancingGPU-driven instancing with compute-updated transforms.
bouncing_linesLINE_LIST topology with simple compute-driven physics.
waveformLINE_STRIP waveform visualizer.

Fragment Shader Effects

Screen-space effects with no geometry beyond a fullscreen triangle.

ExampleWhat it demonstrates
plasmaDemoscene plasma effect.
tunnelFlying-through-a-tunnel polar-coordinate effect.
metaballsMetaball field rendering.
mandelbrotInteractive Mandelbrot explorer.

Interactive and Multi-Window

Input handling, runtime state changes, and more than one surface per device.

ExampleWhat it demonstrates
digital_clockSeven-segment clock display with CPU-generated geometry.
starfield3D starfield with compute-driven star recycling.
particlesRain and snow particle system with a runtime mode switch.
multi_windowThree windows sharing one device.

Shared Code

Examples share a small amount of scaffolding — FPS reporting, run limits, hidden-window creation, headless RGBA capture (GOLDY_EXAMPLE_CAPTURE), and surface-matched pipeline rebuilds. See Shared Helpers.

triangle

The smallest complete Goldy program: three coloured vertices in a retained Scheme, rendered into a scheme-leased offscreen render target, then handed to the swapchain through a SurfaceExchange transaction. Every other windowed example is a variation on this skeleton.

cargo run --features examples --example triangle

What it demonstrates

  • Runtime::acquire_buffer_with_data for a static vertex buffer
  • Scheme::render_pass recorded once and resubmitted every frame
  • SurfaceExchange::bind_render_target plus (&mut submission >> &present).take()? to present
  • Pipeline and scheme rebuild on window resize

Source

examples/triangle.rs:

//! Triangle example - render a colored triangle in an interactive window.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: cargo run --example triangle --features examples

use goldy::{
    shader::builtins, Buffer, BufferKind, Color, Instance, Lease, LeaseRenderTarget, MemoryExchange, NodeAccess,
    RenderPipeline, RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig,
    SurfaceExchange, TargetLoad, Texture, TextureFormat, Transaction, Vertex2D,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::{CaptureDump, FpsWindow};

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    vertex_buffer: Option<Buffer>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    frame_count: u64,
    /// Set after GPU init; FPS excludes startup / shader compile.
    perf_start: Option<Instant>,
    fps_window: FpsWindow,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            vertex_buffer: None,
            pipeline: None,
            shader: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            frame_count: 0,
            perf_start: None,
            fps_window: FpsWindow::new(5.0),
        })
    }

    /// Trailing FPS window in seconds.
    fn fps_window_secs() -> f64 {
        5.0
    }

    /// Soak duration before auto-exit (`GOLDY_EXAMPLE_TIMEOUT` / `EXAMPLE_TIMEOUT` override).
    fn soak_secs() -> f64 {
        common::run_limit_secs().unwrap_or(60.0)
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: Vertex2D::layout(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        vertex_buffer: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
        bg_color: Color,
    ) {
        let mut pass = scheme.render_pass("triangle", scene_rt, TargetLoad::Clear(bg_color));
        pass.with_parcel(vertex_buffer, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, vertex_buffer);
        pass.draw(0..3, 0..1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let vertices = [
            Vertex2D::new(0.0, -0.5, Color::RED),
            Vertex2D::new(-0.5, 0.5, Color::GREEN),
            Vertex2D::new(0.5, 0.5, Color::BLUE),
        ];
        let vertex_buffer = device.acquire_buffer_with_data(&vertices, BufferKind::Scattered)?;

        let shader = ShaderModule::from_slang(&device, builtins::VERTEX_COLOR_2D)?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        let bg_color = Color {
            r: 0.1,
            g: 0.1,
            b: 0.2,
            a: 1.0,
        };
        Self::record_pass(&mut scheme, &pipeline, &vertex_buffer, &scene_rt, bg_color);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.vertex_buffer = Some(vertex_buffer);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.perf_start = Some(Instant::now());
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        self.frame_count += 1;
        if self.perf_start.is_some() {
            self.fps_window.record(Instant::now());
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(vertex_buffer)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.vertex_buffer.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        let bg_color = Color {
                            r: 0.1,
                            g: 0.1,
                            b: 0.2,
                            a: 1.0,
                        };
                        Self::record_pass(&mut scheme, pipeline, vertex_buffer, &rt, bg_color);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let Some(perf_start) = self.perf_start else {
            return;
        };
        let now = Instant::now();
        let elapsed = perf_start.elapsed().as_secs_f64();
        let (window_frames, window_secs, fps) = self.fps_window.stats(now).unwrap_or((0, 0.0, 0.0));
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s last_{:.0}s_fps={fps:.1} (window_frames={window_frames} window_secs={window_secs:.2} present=Auto soak={:.0}s)",
            self.frame_count,
            Self::fps_window_secs(),
            Self::soak_secs()
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Animated Triangle (Scheme + Present)",
                        800,
                        600,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.perf_start.unwrap_or_else(Instant::now));
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Triangle Example (Scheme + Present)");
    println!(
        "PresentMode::Auto (vsync). Auto-exits after {:.0}s soak.",
        App::soak_secs()
    );
    println!(
        "Reports FPS over the last {:.0}s window at exit.",
        App::fps_window_secs()
    );
    println!("Press Escape or close window to exit early.\n");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/vertex_color_2d.slang:

// Simple 2D vertex + fragment shader for colored vertices.
// Used by: triangle, particles, starfield, bouncing_lines, spinning_cube, instancing, waveform

struct VertexInput {
    float2 position : POSITION;
    float4 color : COLOR;
};

struct VertexOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[goldy_vertex]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.color = input.color;
    return output;
}

[goldy_fragment]
float4 fs_main(VertexOutput input) : SV_Target {
    return input.color;
}

mesh_triangle

The same present path as triangle, but the geometry is produced by a [goldy_mesh] entry point driven by dispatch_mesh instead of a vertex buffer. MeshOutput and FsIn deliberately use different struct names so the example also shows Goldy linking stages by semantic (SV_Position, COLOR) rather than by type identity.

cargo run --features examples --example mesh_triangle

What it demonstrates

  • MeshPipeline and dispatch_mesh
  • Automatic payload linking between mesh and fragment stages
  • Capability probing — the example exits 0 when RuntimeCapabilities::mesh_shaders is false

Notes

Mesh shaders are not implemented on the WebGPU backend, and adapters without mesh-shader support skip the example rather than failing.

Source

examples/mesh_triangle.rs:

//! Mesh-shader triangle — `[goldy_mesh]` + `dispatch_mesh` with automatic payload linking.
//!
//! `MeshOutput` and `FsIn` use different struct names; Goldy links them by `SV_Position` / `COLOR`.
//!
//! Skips (exit 0) when `RuntimeCapabilities::mesh_shaders` is false.
//!
//! Run with: cargo run --example mesh_triangle --features examples

use goldy::{
    Color, Instance, Lease, LeaseRenderTarget, MemoryExchange, MeshPipeline, RequestAdapterOptions, RuntimeDescriptor,
    Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat, Transaction,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::{CaptureDump, FpsWindow};

const MESH_SLANG: &str = r#"
import goldy_exp;

struct MeshOutput {
    float4 pos : SV_Position;
    float4 color : COLOR;
};

[goldy_mesh]
[numthreads(1, 1, 1)]
[outputtopology("triangle")]
void mesh_main(out vertices MeshOutput verts[3], out indices uint3 tris[1]) {
    SetMeshOutputCounts(3, 1);
    verts[0] = { float4(0.0, -0.5, 0.0, 1.0), float4(1.0, 0.0, 0.0, 1.0) };
    verts[1] = { float4(-0.5, 0.5, 0.0, 1.0), float4(0.0, 1.0, 0.0, 1.0) };
    verts[2] = { float4(0.5, 0.5, 0.0, 1.0), float4(0.0, 0.0, 1.0, 1.0) };
    tris[0] = uint3(0, 1, 2);
}

struct FsIn {
    float4 pos : SV_Position;
    float4 color : COLOR;
};

[goldy_fragment]
float4 fs_main(FsIn input) : SV_Target {
    return input.color;
}
"#;

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<MeshPipeline>,
    shader: Option<ShaderModule>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    frame_count: u64,
    perf_start: Option<Instant>,
    fps_window: FpsWindow,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            frame_count: 0,
            perf_start: None,
            fps_window: FpsWindow::new(5.0),
        })
    }

    fn fps_window_secs() -> f64 {
        5.0
    }

    fn soak_secs() -> f64 {
        common::run_limit_secs().unwrap_or(60.0)
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<MeshPipeline> {
        MeshPipeline::builder(device)
            .mesh(shader)
            .fragment(shader)
            .target_format(format)
            .build()
    }

    fn record_pass(scheme: &mut Scheme, pipeline: &MeshPipeline, scene_rt: &Lease<LeaseRenderTarget>, bg_color: Color) {
        let mut pass = scheme.render_pass("mesh", scene_rt, TargetLoad::Clear(bg_color));
        pass.set_mesh_pipeline(pipeline);
        pass.dispatch_mesh(1, 1, 1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        if !device.capabilities().mesh_shaders {
            println!("skip: RuntimeCapabilities::mesh_shaders is false on this adapter");
            std::process::exit(0);
        }
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, MESH_SLANG)?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        let bg_color = Color {
            r: 0.1,
            g: 0.1,
            b: 0.2,
            a: 1.0,
        };
        Self::record_pass(&mut scheme, &pipeline, &scene_rt, bg_color);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.perf_start = Some(Instant::now());
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }
        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        self.frame_count += 1;
        if self.perf_start.is_some() {
            self.fps_window.record(Instant::now());
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader)) = (self.ctx.as_ref(), self.device.as_ref(), self.shader.as_ref())
        {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        let bg_color = Color {
                            r: 0.1,
                            g: 0.1,
                            b: 0.2,
                            a: 1.0,
                        };
                        Self::record_pass(&mut scheme, pipeline, &rt, bg_color);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let Some(perf_start) = self.perf_start else {
            return;
        };
        let now = Instant::now();
        let elapsed = perf_start.elapsed().as_secs_f64();
        let (window_frames, window_secs, fps) = self.fps_window.stats(now).unwrap_or((0, 0.0, 0.0));
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s last_{:.0}s_fps={fps:.1} (window_frames={window_frames} window_secs={window_secs:.2} present=Auto soak={:.0}s)",
            self.frame_count,
            Self::fps_window_secs(),
            Self::soak_secs()
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Mesh Triangle", 800, 600))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.perf_start.unwrap_or_else(Instant::now));
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Mesh Triangle (set_mesh_pipeline + dispatch_mesh)");
    println!("Press Escape or close window to exit.\n");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

The Slang source is inline in the example above.

gradient

An animated full-screen gradient driven by a single time uniform. Rendering is vertex-less: the vertex stage synthesizes a fullscreen triangle from SV_VertexID, which is the Goldy-native way to write screen-space effects.

cargo run --features examples --example gradient

What it demonstrates

  • Vertex-less fullscreen rendering
  • Per-frame uniform updates through a deposit transaction
  • LayoutCheckable host/shader struct layout validation

Notes

Set GOLDY_VALIDATE_LAYOUTS=1 to have the example cross-check its uniform struct layout against Slang reflection:

GOLDY_VALIDATE_LAYOUTS=1 cargo run --features examples --example gradient

Source

examples/gradient.rs:

//! Gradient example - animated color gradient.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//! Uses vertex-less fullscreen triangle (Goldy-native pattern).
//!
//! Run with: `cargo run --example gradient`
//!

use goldy::{
    shaders, Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct TimeUniforms {
    time: f32,
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    uniform: Option<Buffer>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    uniform_deposit: Option<DepositTransaction>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            uniform: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            uniform_deposit: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        uniform: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("gradient", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(uniform, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.draw_fullscreen();
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang_with_gpu_types(&device, shaders::GRADIENT, &[TimeUniforms::GPU_TYPE])?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let uniform = device.acquire_buffer_sized::<TimeUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &uniform, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<TimeUniforms>(&uniform, 1),
        )?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.uniform = Some(uniform);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.upload_scheme = Some(upload_scheme);
        self.uniform_deposit = Some(uniform_deposit);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());
        let uniforms = TimeUniforms { time };
        let upload = self.upload_scheme.as_mut().unwrap();
        (self.uniform_deposit.as_ref().unwrap() << &uniforms)?;
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(uniform)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.uniform.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, uniform, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Animated Gradient (Scheme + Present)",
                        800,
                        600,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Gradient Example - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/gradient.slang:

// Animated color gradient
// Uses vertex-less fullscreen triangle and time uniform

import goldy_exp;


[goldy_vertex]
FullscreenVarying vs_main(VertexId vertex_id) {
    return vs_fullscreen_triangle(vertex_id.value);
}

[goldy_fragment]
float4 fs_main(TimeUniforms uniforms, FullscreenVarying input) : SV_Target {
    float2 uv = input.uv;
    float t = uniforms.time;
    
    // Smooth animated gradient using sine waves
    float3 color;
    color.r = 0.5 + 0.5 * sin(uv.x * 3.14159 + t * 0.7);
    color.g = 0.5 + 0.5 * sin(uv.y * 3.14159 + t * 0.5 + 2.0);
    color.b = 0.5 + 0.5 * sin((uv.x + uv.y) * 2.0 + t * 0.3 + 4.0);
    
    // Add subtle variation based on position
    color *= 0.8 + 0.2 * sin(uv.x * 6.28 + uv.y * 6.28 + t);
    
    return float4(color, 1.0);
}

checkerboard

A procedural checkerboard whose UVs are distorted over time in the fragment shader — no texture and no vertex buffer, just a time uniform and arithmetic.

cargo run --features examples --example checkerboard

What it demonstrates

  • Procedural fragment-shader texturing
  • ShaderModule::from_slang_with_options for compile-time options
  • Retained scheme with offscreen render pass and copy-to-present

Notes

Set GOLDY_VALIDATE_LAYOUTS=1 to validate the uniform layout against Slang reflection.

Source

examples/checkerboard.rs:

//! Checkerboard example - procedural texture with animation.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example checkerboard`
//!

use goldy::{
    shaders, Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct TimeUniforms {
    time: f32,
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    uniform: Option<Buffer>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    uniform_deposit: Option<DepositTransaction>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            uniform: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            uniform_deposit: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        uniform: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("checkerboard", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(uniform, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.draw_fullscreen();
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader =
            ShaderModule::from_slang_with_gpu_types(&device, shaders::CHECKERBOARD, &[TimeUniforms::GPU_TYPE])?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let uniform = device.acquire_buffer_sized::<TimeUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &uniform, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<TimeUniforms>(&uniform, 1),
        )?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.uniform = Some(uniform);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.upload_scheme = Some(upload_scheme);
        self.uniform_deposit = Some(uniform_deposit);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());
        let uniforms = TimeUniforms { time };
        let upload = self.upload_scheme.as_mut().unwrap();
        (self.uniform_deposit.as_ref().unwrap() << &uniforms)?;
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(uniform)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.uniform.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, uniform, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Animated Checkerboard (Scheme + Present)",
                        800,
                        800,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Checkerboard Example - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/checkerboard.slang:

// Animated checkerboard pattern
// Uses vertex-less fullscreen triangle and time uniform

import goldy_exp;


[goldy_vertex]
FullscreenVarying vs_main(VertexId vertex_id) {
    return vs_fullscreen_triangle(vertex_id.value);
}

[goldy_fragment]
float4 fs_main(TimeUniforms uniforms, FullscreenVarying input) : SV_Target {
    float t = uniforms.time;
    
    // Animated wave distortion
    float2 uv = input.uv;
    uv.x += sin(uv.y * 10.0 + t * 2.0) * 0.02;
    uv.y += cos(uv.x * 10.0 + t * 1.5) * 0.02;
    
    // Scale for checker pattern
    float scale = 8.0;
    float2 checker = floor(uv * scale);
    bool is_white = fmod(checker.x + checker.y, 2.0) == 0.0;
    
    // Animate colors
    float3 color1 = float3(
        0.2 + 0.1 * sin(t),
        0.1 + 0.1 * cos(t * 1.3),
        0.3 + 0.1 * sin(t * 0.7)
    );
    float3 color2 = float3(
        0.9 + 0.1 * cos(t * 0.8),
        0.85 + 0.1 * sin(t * 1.1),
        0.8 + 0.1 * cos(t)
    );
    
    // Add subtle gradient
    color1 += uv.y * 0.1;
    color2 -= uv.y * 0.1;
    
    if (is_white) {
        return float4(color2, 1.0);
    } else {
        return float4(color1, 1.0);
    }
}

compute_particles

A compute shader integrates particle positions in place, and a graphics pass draws them as instanced quads from the same buffer. Both nodes live in one retained scheme, so Goldy derives the compute-to-raster barrier from the declared parcel accesses.

cargo run --features examples --example compute_particles

What it demonstrates

  • Compute and render nodes in a single retained scheme
  • Read/write parcel access driving automatic hazard tracking
  • Instanced draws sourced from compute output

Source

examples/compute_particles.rs:

//! GPU Particle Simulation Example
//!
//! Demonstrates retained scheme with compute dispatch → offscreen render → copy-to-present.
//!
//! Run with: `cargo run --example compute_particles`

use anyhow::Result;
use goldy::{
    Buffer, BufferFlags, BufferKind, Color, ComputePipeline, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc,
    RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad,
    Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const NUM_PARTICLES: u32 = 1024;

#[goldy::gpu]
struct Particle {
    position: [f32; 2],
    velocity: [f32; 2],
}

#[goldy::gpu]
struct SimParams {
    delta_time: f32,
}

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut state = RenderState::new(None)?;
        while !state.capture_done() {
            state.render()?;
        }
        return Ok(());
    }

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(winit::event_loop::ControlFlow::Poll);

    let mut app = App::default();
    event_loop.run_app(&mut app)?;

    Ok(())
}

#[derive(Default)]
struct App {
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    device: Arc<goldy::Runtime>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    present: Option<Transaction>,
    scheme: Scheme,
    scene_rt: Lease<LeaseRenderTarget>,
    compute_pipeline: ComputePipeline,
    render_shader: ShaderModule,
    render_pipeline: RenderPipeline,
    particle_buffer: Buffer,
    params_buffer: Buffer,
    upload_scheme: Scheme,
    params_deposit: DepositTransaction,
    frame_count: u32,
    start_time: std::time::Instant,
    last_frame_time: std::time::Instant,
}

impl RenderState {
    fn create_render_pipeline(
        device: &goldy::Runtime,
        render_shader: &ShaderModule,
        format: TextureFormat,
    ) -> Result<RenderPipeline> {
        common::render_pipeline(
            device,
            render_shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                topology: PrimitiveTopology::TriangleList,
                ..Default::default()
            },
        )
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn record_scheme(
        scheme: &mut Scheme,
        compute_pipeline: &ComputePipeline,
        render_pipeline: &RenderPipeline,
        particle_buffer: &Buffer,
        params_buffer: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        scheme
            .node("update_particles", compute_pipeline)
            .with_parcel(particle_buffer, NodeAccess::ReadWrite)
            .with_parcel(params_buffer, NodeAccess::Read)
            .dispatch(NUM_PARTICLES.div_ceil(64), 1, 1);

        let bg_color = Color {
            r: 0.03,
            g: 0.02,
            b: 0.08,
            a: 1.0,
        };

        let mut pass = scheme.render_pass("particles", scene_rt, TargetLoad::Clear(bg_color));
        pass.with_parcel(particle_buffer, NodeAccess::Read);
        pass.set_pipeline(render_pipeline);
        pass.draw(0..6, 0..NUM_PARTICLES);
        pass.finish();
    }

    fn target(&self) -> (TextureFormat, u32, u32) {
        if let Some(surface) = &self.surface {
            let (width, height) = surface.size();
            (surface.format(), width, height)
        } else {
            let capture = self.capture.as_ref().expect("capture dump");
            let (width, height) = capture.size();
            (CaptureDump::format(), width, height)
        }
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn rerecord_scheme(&mut self) {
        let mut scheme = Scheme::new(&self.ctx);
        let (format, width, height) = self.target();
        if let Ok(rt) = self.ctx.lease_render_target(width.max(1), height.max(1), format, None) {
            self.scene_rt = rt;
            Self::record_scheme(
                &mut scheme,
                &self.compute_pipeline,
                &self.render_pipeline,
                &self.particle_buffer,
                &self.params_buffer,
                &self.scene_rt,
            );
            if let Ok(present) = Self::bind_frame(
                &mut scheme,
                &self.scene_rt,
                self.surface.as_ref(),
                self.readback.as_ref(),
            ) {
                self.present = present;
                self.scheme = scheme;
            }
        }
    }

    fn new(window: Option<Arc<Window>>) -> Result<Self> {
        let instance = Instance::new()?;
        let device = Arc::new(
            instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let compute_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/particle_update.slang"),
            &[Particle::GPU_TYPE, SimParams::GPU_TYPE],
        )?;
        let render_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/particle_render.slang"),
            &[Particle::GPU_TYPE],
        )?;

        let mut particles = Vec::with_capacity(NUM_PARTICLES as usize);
        for i in 0..NUM_PARTICLES {
            let t = i as f32 / NUM_PARTICLES as f32;
            let angle = t * std::f32::consts::TAU * 5.0;
            let radius = 0.1 + t * 0.6;

            let noise_x = ((i * 17) % 100) as f32 / 100.0 - 0.5;
            let noise_y = ((i * 31) % 100) as f32 / 100.0 - 0.5;

            particles.push(Particle {
                position: [
                    radius * angle.cos() + noise_x * 0.1,
                    radius * angle.sin() + noise_y * 0.1,
                ],
                velocity: [angle.sin() * 0.3 + noise_x * 0.2, -angle.cos() * 0.3 + noise_y * 0.2],
            });
        }

        let particle_buffer = device.acquire_buffer_with_data(&particles, BufferKind::Scattered)?;
        let params_buffer = device.acquire_buffer_sized::<SimParams>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let compute_pipeline = ComputePipeline::new(&device, &compute_shader)?;
        let render_pipeline = Self::create_render_pipeline(&device, &render_shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_scheme(
            &mut scheme,
            &compute_pipeline,
            &render_pipeline,
            &particle_buffer,
            &params_buffer,
            &scene_rt,
        );
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let params_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<SimParams>(&params_buffer, 1),
        )?;

        println!("Created compute particles example with {NUM_PARTICLES} particles (Scheme + Present)");

        Ok(Self {
            window,
            device,
            ctx,
            surface,
            capture,
            readback,
            present,
            scheme,
            scene_rt,
            compute_pipeline,
            render_shader,
            render_pipeline,
            particle_buffer,
            params_buffer,
            upload_scheme,
            params_deposit,
            frame_count: 0,
            start_time: std::time::Instant::now(),
            last_frame_time: std::time::Instant::now(),
        })
    }

    fn render(&mut self) -> Result<()> {
        self.frame_count += 1;

        let dt = if let Some(capture) = &self.capture {
            capture.dt()
        } else {
            self.last_frame_time.elapsed().as_secs_f32()
        }
        .min(0.05);
        self.last_frame_time = std::time::Instant::now();

        (&self.params_deposit << &SimParams { delta_time: dt })?;
        self.upload_scheme.submit()?;

        let mut submission = self.scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        if let Some(window) = &self.window {
            window.request_redraw();
        }
        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Compute Particles", 800, 600))
                    .expect("Failed to create window"),
            );

            match RenderState::new(Some(window.clone())) {
                Ok(mut state) => {
                    if let Err(e) = state.render() {
                        tracing::error!("First frame error: {e}");
                    }
                    common::reveal_window(&window);
                    self.state = Some(state);
                    window.request_redraw();
                }
                Err(e) => {
                    tracing::error!("Failed to create render state: {e:?}");
                    event_loop.exit();
                }
            }
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::Resized(size) => {
                if let Some(state) = &mut self.state {
                    if size.width > 0 && size.height > 0 {
                        let Some(surface) = state.surface.as_ref() else {
                            return;
                        };
                        let (prev_w, prev_h) = surface.size();
                        if size.width == prev_w && size.height == prev_h {
                            return;
                        }
                        let _ = surface.resize(size.width, size.height);
                        let format = surface.format();
                        if let Ok(pipeline) =
                            RenderState::create_render_pipeline(&state.device, &state.render_shader, format)
                        {
                            state.render_pipeline = pipeline;
                        }
                        state.rerecord_scheme();
                    }
                }
            }
            WindowEvent::RedrawRequested => {
                if let Some(state) = &mut self.state {
                    if let Err(e) = state.render() {
                        tracing::error!("Render error: {e}");
                    }
                }
            }
            _ => {}
        }
    }
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/particle_update.slang:

// Particle simulation compute shader
// Updates particle positions and velocities

import goldy_exp;


static const float2 gravity = float2(0.0, -0.3);
// Per-second damping factor (0.998^60 ≈ 0.887 — tuned at 60fps originally)
static const float dampingPerSecond = 0.887;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Particle> PARTICLES, SimParams params, ThreadId id) {
    uint idx = id.x;
    
    if (idx >= 65536) return;
    
    Particle p = PARTICLES[idx];
    
    if (p.position.x == 0.0 && p.position.y == 0.0 && 
        p.velocity.x == 0.0 && p.velocity.y == 0.0) {
        return;
    }
    
    p.velocity += gravity * params.delta_time;
    p.velocity *= pow(dampingPerSecond, params.delta_time);
    p.position += p.velocity * params.delta_time;
    
    if (p.position.x < -0.95) {
        p.velocity.x = abs(p.velocity.x) * 0.8;
        p.position.x = -0.95;
    }
    if (p.position.x > 0.95) {
        p.velocity.x = -abs(p.velocity.x) * 0.8;
        p.position.x = 0.95;
    }
    if (p.position.y < -0.95) {
        p.velocity.y = abs(p.velocity.y) * 0.8;
        p.position.y = -0.95;
    }
    if (p.position.y > 0.95) {
        p.velocity.y = -abs(p.velocity.y) * 0.8;
        p.position.y = 0.95;
    }
    
    PARTICLES[idx] = p;
}

shaders/particle_render.slang:

// Particle rendering shader
// Visualizes particles as colored quads using instancing

import goldy_exp;


struct VSOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

static const float2 quadVerts[6] = {
    float2(-1, -1), float2( 1, -1), float2( 1,  1),
    float2(-1, -1), float2( 1,  1), float2(-1,  1)
};

[goldy_vertex]
VSOutput vs_main(Scattered<Particle> PARTICLES, VertexId vertexID, InstanceId instanceID) {
    VSOutput output;
    
    Particle p = PARTICLES[instanceID.value];
    
    float size = 0.015;
    float2 localPos = quadVerts[vertexID.value] * size;
    
    output.position = float4(p.position + localPos, 0.0, 1.0);
    
    float speed = length(p.velocity);
    float angle = atan2(p.velocity.y, p.velocity.x);
    
    float h = (angle + 3.14159) / (2.0 * 3.14159);
    float3 rgb = abs(h * 6.0 - float3(3, 2, 4)) * float3(1, -1, -1) + float3(-1, 2, 2);
    rgb = clamp(rgb, 0.0, 1.0);
    
    float brightness = clamp(speed * 2.0 + 0.3, 0.3, 1.0);
    output.color = float4(rgb * brightness, 1.0);
    
    return output;
}

[goldy_fragment]
float4 fs_main(VSOutput input) : SV_Target {
    return input.color;
}

game_of_life

Conway's Game of Life on the GPU. Both cell grids live in a single retained record buffer as fields "a" and "b", so ping-pong is a sub-view swap rather than two separate parcels. Two schemes are recorded once (AB writes b, BA writes a) and resubmitted on alternate steps. Idle redraws skip submit and leave the last present on the surface.

cargo run --features examples --example game_of_life

What it demonstrates

  • Sub-views of one retained mosaic parcel for ping-pong state
  • Two retained schemes, alternating resubmit
  • Compute → render → present in each scheme
  • Skipping submit on idle redraws

Source

examples/game_of_life.rs:

//! Conway's Game of Life — two retained schemes, alternating resubmit.
//!
//! Ping-pong cell grids live in one retained record buffer (fields `"a"` / `"b"`).
//! Orientation AB reads `a` and writes `b`; BA is the swap. Each is recorded once
//! (and again on resize). Simulation steps alternate which scheme submits. Idle
//! redraws skip submit and leave the last present on the surface.
//!
//! Run with: `cargo run --example game_of_life`

use anyhow::Result;
use goldy::{
    field, Buffer, ComputePipeline, Context, Init, Instance, Lease, LeaseRenderTarget, NodeAccess, PrimitiveTopology,
    RenderPipeline, RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, Submission,
    SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const GRID_WIDTH: u32 = 128;
const GRID_HEIGHT: u32 = 128;
const CELL_COUNT: u32 = GRID_WIDTH * GRID_HEIGHT;

fn record_scheme(
    scheme: &mut Scheme,
    cells: &Buffer,
    read_field: &str,
    write_field: &str,
    compute_pipeline: &ComputePipeline,
    render_pipeline: &RenderPipeline,
    scene_rt: &Lease<LeaseRenderTarget>,
) {
    scheme
        .node("game_of_life", compute_pipeline)
        .with_parcel(&cells[read_field], NodeAccess::Read)
        .with_parcel(&cells[write_field], NodeAccess::Overwrite)
        .dispatch(GRID_WIDTH.div_ceil(8), GRID_HEIGHT.div_ceil(8), 1);

    let mut pass = scheme.render_pass("game_of_life_render", scene_rt, TargetLoad::Discard);
    pass.with_parcel(&cells[write_field], NodeAccess::Read);
    pass.set_pipeline(render_pipeline);
    pass.draw(0..3, 0..1);
    pass.finish();
}

struct FrameBind<'a> {
    format: TextureFormat,
    width: u32,
    height: u32,
    surface: Option<&'a SurfaceExchange>,
    readback: Option<&'a Texture>,
}

struct Recorded {
    scheme: Scheme,
    present: Option<Transaction>,
}

fn bind_frame(
    scheme: &mut Scheme,
    scene_rt: &Lease<LeaseRenderTarget>,
    bind: &FrameBind<'_>,
) -> anyhow::Result<Option<Transaction>> {
    if let Some(surface) = bind.surface {
        let present = surface.bind_render_target(scheme, scene_rt)?;
        Ok(Some(present))
    } else {
        let readback = bind.readback.expect("capture readback");
        scheme.copy_to_texture(scene_rt, readback)?;

        Ok(None)
    }
}

fn build_scheme(
    ctx: &Context,
    cells: &Buffer,
    read_field: &str,
    write_field: &str,
    compute_pipeline: &ComputePipeline,
    render_pipeline: &RenderPipeline,
    bind: &FrameBind<'_>,
) -> anyhow::Result<Recorded> {
    let mut scheme = Scheme::new(ctx);
    let scene_rt = ctx.lease_render_target(bind.width.max(1), bind.height.max(1), bind.format, None)?;
    record_scheme(
        &mut scheme,
        cells,
        read_field,
        write_field,
        compute_pipeline,
        render_pipeline,
        &scene_rt,
    );
    let present = bind_frame(&mut scheme, &scene_rt, bind)?;
    Ok(Recorded { scheme, present })
}

fn build_schemes(
    ctx: &Context,
    cells: &Buffer,
    compute_pipeline: &ComputePipeline,
    render_pipeline: &RenderPipeline,
    bind: &FrameBind<'_>,
) -> anyhow::Result<(Recorded, Recorded)> {
    Ok((
        build_scheme(ctx, cells, "a", "b", compute_pipeline, render_pipeline, bind)?,
        build_scheme(ctx, cells, "b", "a", compute_pipeline, render_pipeline, bind)?,
    ))
}

fn create_initial_state() -> Vec<u32> {
    let mut cells = vec![0u32; CELL_COUNT as usize];

    let gun = [
        (1, 5),
        (1, 6),
        (2, 5),
        (2, 6),
        (11, 5),
        (11, 6),
        (11, 7),
        (12, 4),
        (12, 8),
        (13, 3),
        (13, 9),
        (14, 3),
        (14, 9),
        (15, 6),
        (16, 4),
        (16, 8),
        (17, 5),
        (17, 6),
        (17, 7),
        (18, 6),
        (21, 3),
        (21, 4),
        (21, 5),
        (22, 3),
        (22, 4),
        (22, 5),
        (23, 2),
        (23, 6),
        (25, 1),
        (25, 2),
        (25, 6),
        (25, 7),
        (35, 3),
        (35, 4),
        (36, 3),
        (36, 4),
    ];

    let offset_x = 10;
    let offset_y = 10;
    for (x, y) in gun.iter() {
        let px = (x + offset_x) as u32;
        let py = (y + offset_y) as u32;
        if px < GRID_WIDTH && py < GRID_HEIGHT {
            cells[(py * GRID_WIDTH + px) as usize] = 1;
        }
    }

    let mut rng = 42u64;
    for y in 60..100 {
        for x in 60..100 {
            rng = rng.wrapping_mul(6364136223846793005).wrapping_add(1);
            if (rng >> 32).is_multiple_of(4) {
                cells[(y * GRID_WIDTH + x) as usize] = 1;
            }
        }
    }

    cells
}

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut state = RenderState::new(None)?;
        while !state.capture_done() {
            state.render()?;
        }
        return Ok(());
    }

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App::default();
    event_loop.run_app(&mut app)?;

    Ok(())
}

#[derive(Default)]
struct App {
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    ctx: Context,
    surface: Option<SurfaceExchange>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scheme_ab: Scheme,
    scheme_ba: Scheme,
    present_ab: Option<Transaction>,
    present_ba: Option<Transaction>,
    compute_pipeline: ComputePipeline,
    render_pipeline: RenderPipeline,
    cells: Buffer,
    use_buffer_a: bool,
    frame_count: u32,
    last_update: std::time::Instant,
    start_time: std::time::Instant,
}

impl RenderState {
    fn target(&self) -> (TextureFormat, u32, u32) {
        if let Some(surface) = &self.surface {
            let (width, height) = surface.size();
            (surface.format(), width, height)
        } else {
            let capture = self.capture.as_ref().expect("capture dump");
            let (width, height) = capture.size();
            (CaptureDump::format(), width, height)
        }
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn record_orientations(&mut self) -> Result<()> {
        let (format, width, height) = self.target();
        let bind = FrameBind {
            format,
            width,
            height,
            surface: self.surface.as_ref(),
            readback: self.readback.as_ref(),
        };
        let (ab, ba) = build_schemes(
            &self.ctx,
            &self.cells,
            &self.compute_pipeline,
            &self.render_pipeline,
            &bind,
        )?;
        self.scheme_ab = ab.scheme;
        self.present_ab = ab.present;
        self.scheme_ba = ba.scheme;
        self.present_ba = ba.present;
        Ok(())
    }

    fn new(window: Option<Arc<Window>>) -> Result<Self> {
        let instance = Instance::new()?;
        let device = Arc::new(
            instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let compute_shader = ShaderModule::from_slang(&device, include_str!("../shaders/game_of_life.slang"))?;
        let render_shader = ShaderModule::from_slang(&device, include_str!("../shaders/game_of_life_render.slang"))?;

        let initial_state = create_initial_state();
        let cells = device.acquire_record([
            field("a", Init::data(&initial_state)),
            field("b", Init::data(&initial_state)),
        ])?;

        let compute_pipeline = ComputePipeline::new(&device, &compute_shader)?;
        let render_pipeline = RenderPipeline::new(
            &device,
            &render_shader,
            &render_shader,
            &RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::default(),
                topology: PrimitiveTopology::TriangleList,
                target_format: format,
                ..Default::default()
            },
        )?;

        let bind = FrameBind {
            format,
            width,
            height,
            surface: surface.as_ref(),
            readback: readback.as_ref(),
        };
        let (ab, ba) = build_schemes(&ctx, &cells, &compute_pipeline, &render_pipeline, &bind)?;

        println!("Game of Life initialized: {}x{} grid", GRID_WIDTH, GRID_HEIGHT);
        println!("Features Gosper Glider Gun + random cells");
        println!("Press Escape or close window to exit");

        Ok(Self {
            window,
            ctx,
            surface,
            capture,
            readback,
            scheme_ab: ab.scheme,
            scheme_ba: ba.scheme,
            present_ab: ab.present,
            present_ba: ba.present,
            compute_pipeline,
            render_pipeline,
            cells,
            use_buffer_a: true,
            frame_count: 0,
            last_update: std::time::Instant::now(),
            start_time: std::time::Instant::now(),
        })
    }

    fn settle(
        present: Option<&Transaction>,
        readback: Option<&Texture>,
        capture: Option<&mut CaptureDump>,
        submission: &mut Submission,
    ) -> Result<()> {
        if let Some(present) = present {
            (submission >> present).take()?;
        } else {
            let pixels = (submission >> readback.expect("capture readback"))
                .take::<u8>()?
                .to_vec();
            capture.expect("capture dump").write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn step(&mut self) -> Result<()> {
        if self.use_buffer_a {
            let mut submission = self.scheme_ab.submit()?;
            Self::settle(
                self.present_ab.as_ref(),
                self.readback.as_ref(),
                self.capture.as_mut(),
                &mut submission,
            )?;
        } else {
            let mut submission = self.scheme_ba.submit()?;
            Self::settle(
                self.present_ba.as_ref(),
                self.readback.as_ref(),
                self.capture.as_mut(),
                &mut submission,
            )?;
        }
        self.use_buffer_a = !self.use_buffer_a;
        Ok(())
    }

    fn render(&mut self) -> Result<()> {
        let now = std::time::Instant::now();
        let should_step =
            self.capture.is_some() || self.frame_count == 0 || now.duration_since(self.last_update).as_millis() > 33;

        if should_step {
            self.last_update = now;
            self.step()?;
            self.frame_count += 1;
        }

        if let Some(window) = &self.window {
            window.request_redraw();
        }
        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Game of Life", 800, 800))
                    .expect("Failed to create window"),
            );

            match RenderState::new(Some(window.clone())) {
                Ok(mut state) => {
                    if let Err(e) = state.render() {
                        tracing::error!("First frame error: {e}");
                    }
                    common::reveal_window(&window);
                    self.state = Some(state);
                    window.request_redraw();
                }
                Err(e) => {
                    tracing::error!("Failed to create render state: {:#}", e);
                    event_loop.exit();
                }
            }
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => {
                event_loop.exit();
            }
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::Resized(size) => {
                if let Some(state) = &mut self.state {
                    if size.width > 0 && size.height > 0 {
                        let Some(surface) = state.surface.as_ref() else {
                            return;
                        };
                        if let Err(e) = surface.resize(size.width, size.height) {
                            tracing::error!("Failed to resize surface: {e}");
                            return;
                        }
                        if let Err(e) = state.record_orientations() {
                            tracing::error!("Failed to rerecord schemes: {e}");
                            return;
                        }
                        state.last_update = std::time::Instant::now()
                            .checked_sub(std::time::Duration::from_millis(34))
                            .unwrap_or(state.last_update);
                    }
                }
            }
            WindowEvent::RedrawRequested => {
                if let Some(state) = &mut self.state {
                    if let Err(e) = state.render() {
                        tracing::error!("Render error: {:#}", e);
                    }
                }
            }
            _ => {}
        }
    }
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/game_of_life.slang:

// Conway's Game of Life compute shader
// Uses ping-pong buffers: reads from one, writes to the other

import goldy_exp;

static const uint GRID_WIDTH = 128;
static const uint GRID_HEIGHT = 128;

uint getCell(Scattered<uint> state, int x, int y) {
    x = (x + GRID_WIDTH) % GRID_WIDTH;
    y = (y + GRID_HEIGHT) % GRID_HEIGHT;
    return state[y * GRID_WIDTH + x];
}

uint countNeighbors(Scattered<uint> state, int x, int y) {
    uint count = 0;
    count += getCell(state, x - 1, y - 1);
    count += getCell(state, x,     y - 1);
    count += getCell(state, x + 1, y - 1);
    count += getCell(state, x - 1, y);
    count += getCell(state, x + 1, y);
    count += getCell(state, x - 1, y + 1);
    count += getCell(state, x,     y + 1);
    count += getCell(state, x + 1, y + 1);
    return count;
}

[goldy_compute]
[numthreads(8, 8, 1)]
void cs_main(Scattered<uint> CURRENT_STATE, Scattered<uint> NEXT_STATE,
             ThreadId id) {
    if (id.x >= GRID_WIDTH || id.y >= GRID_HEIGHT) return;
    
    uint idx = id.y * GRID_WIDTH + id.x;
    uint cell = CURRENT_STATE[idx];
    uint neighbors = countNeighbors(CURRENT_STATE, int(id.x), int(id.y));
    
    // Conway's rules
    uint newState = 0;
    if (cell == 1) {
        newState = (neighbors == 2 || neighbors == 3) ? 1 : 0;
    } else {
        newState = (neighbors == 3) ? 1 : 0;
    }
    
    NEXT_STATE[idx] = newState;
}

shaders/game_of_life_render.slang:

// Game of Life rendering shader
// Renders the grid as a fullscreen quad with cell colors

import goldy_exp;

static const uint GRID_WIDTH = 128;
static const uint GRID_HEIGHT = 128;

struct VSOutput {
    float4 position : SV_Position;
    float2 uv : TEXCOORD0;
};

static const float2 positions[3] = {
    float2(-1, -1),
    float2( 3, -1),
    float2(-1,  3)
};

static const float2 uvs[3] = {
    float2(0, 1),
    float2(2, 1),
    float2(0, -1)
};

[goldy_vertex]
VSOutput vs_main(VertexId vertexID) {
    VSOutput output;
    output.position = float4(positions[vertexID.value], 0.0, 1.0);
    output.uv = uvs[vertexID.value];
    return output;
}

[goldy_fragment]
float4 fs_main(Scattered<uint> CELLS, VSOutput input) : SV_Target {
    float2 uv = input.uv;
    uv.y = 1.0 - uv.y;
    
    int x = int(uv.x * GRID_WIDTH);
    int y = int(uv.y * GRID_HEIGHT);
    
    x = clamp(x, 0, int(GRID_WIDTH) - 1);
    y = clamp(y, 0, int(GRID_HEIGHT) - 1);
    
    uint idx = y * GRID_WIDTH + x;
    uint cell = CELLS[idx];

    float2 cellUV = frac(float2(uv.x * GRID_WIDTH, uv.y * GRID_HEIGHT));
    float gridLine = (cellUV.x < 0.05 || cellUV.y < 0.05) ? 0.15 : 0.0;
    
    if (cell == 1) {
        float3 alive = float3(0.2, 0.9, 0.3);
        return float4(alive + gridLine, 1.0);
    } else {
        float3 dead = float3(0.05, 0.08, 0.1);
        return float4(dead + gridLine, 1.0);
    }
}

compute_to_surface

Rendering with no RenderPipeline at all. The compute shader writes the swapchain drawable obtained from SurfaceExchange::bind_destination, and the frame is settled by claiming the present transaction. This is the shortest path from a dispatch to the screen.

cargo run --features examples --example compute_to_surface

What it demonstrates

  • SurfaceExchange::bind_destination — present-on-scheme
  • Transaction::claim and Claim::consume settlement
  • gpu::DirectSpatial<gpu::Float4> storage-texture writes to a drawable

Notes

On the WebGPU backend the swapchain image cannot be bound as storage, so present falls back to a copy or blit path automatically. See Backend Architecture for the GOLDY_WEBGPU_PRESENT override.

Source

examples/compute_to_surface.rs:

//! Compute-to-Surface example — pure compute rendering without a graphics pipeline.
//!
//! Demonstrates present-on-scheme: a retained [`Scheme`] writes directly to a
//! drawable from [`SurfaceExchange::bind_destination`], then presents via
//! [`Transaction::claim`] and [`Claim::consume`].
//!
//! Run with: cargo run --example compute_to_surface

use anyhow::Result;
use goldy::{
    Buffer, BufferKind, DepositTarget, DepositTransaction, Instance, MemoryExchange, PresentMode,
    RequestAdapterOptions, RuntimeDescriptor, Scheme, SurfaceConfig, SurfaceExchange, Texture, Transaction,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct Uniforms {
    width: u32,
    height: u32,
    time: f32,
}

#[goldy::compute(workgroup_size = [8, 8, 1])]
fn plasma(uniforms: &[Uniforms], output: goldy::gpu::DirectSpatial<goldy::gpu::Float4>) {
    let tid = goldy::gpu::global_id();
    let u: Uniforms = uniforms[0];
    if tid.x >= u.width || tid.y >= u.height {
        return;
    }

    let uv = goldy::gpu::float2(tid.x as f32 / u.width as f32, tid.y as f32 / u.height as f32);
    let mut p = uv * 2.0 - 1.0;
    p.x *= u.width as f32 / u.height as f32;

    let mut v = 0.0;
    v += goldy::gpu::sin(p.x * 6.0 + u.time);
    v += goldy::gpu::sin(p.y * 6.0 + u.time * 1.3);
    v += goldy::gpu::sin((p.x + p.y) * 4.0 + u.time * 0.7);
    v += goldy::gpu::sin(goldy::gpu::length(p) * 8.0 - u.time * 2.0);
    v *= 0.25;

    let col = goldy::gpu::float3(
        0.5 + 0.5 * goldy::gpu::sin(v * 3.14159 + 0.0),
        0.5 + 0.5 * goldy::gpu::sin(v * 3.14159 + 2.094),
        0.5 + 0.5 * goldy::gpu::sin(v * 3.14159 + 4.188),
    );
    output[tid.xy] = goldy::gpu::float4(col.x, col.y, col.z, 1.0);
}

const INITIAL_WIDTH: u32 = 800;
const INITIAL_HEIGHT: u32 = 600;

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let warmup = warm_gpu()?;
        let mut app = App {
            warmup: Some(warmup),
            state: None,
        };
        app.init(None)?;
        let state = app.state.as_mut().expect("capture state");
        while !state.capture.as_ref().is_none_or(CaptureDump::finished) {
            render_frame(state)?;
        }
        return Ok(());
    }

    println!("Goldy — Compute to Surface Example");
    println!("===================================");
    println!("Press V to toggle vsync, Escape to exit\n");

    println!("Initializing GPU...");
    let warmup = warm_gpu()?;
    println!("GPU ready.");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App {
        warmup: Some(warmup),
        state: None,
    };
    event_loop.run_app(&mut app)?;

    Ok(())
}

/// Runtime, context, and compiled compute pipeline — everything except the window/surface.
struct GpuWarmup {
    ctx: goldy::Context,
    kernel: plasma::Kernel,
    device: Arc<goldy::Runtime>,
}

fn warm_gpu() -> Result<GpuWarmup> {
    let instance = Instance::new()?;
    let device = Arc::new(
        instance
            .request_adapter(&RequestAdapterOptions::default())?
            .request_runtime(&RuntimeDescriptor::default())?,
    );
    let ctx = device.create_context()?;
    let kernel = plasma::Kernel::prepare(&device)?;
    Ok(GpuWarmup { ctx, kernel, device })
}

#[derive(Default)]
struct App {
    warmup: Option<GpuWarmup>,
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scheme: Scheme,
    compute_pipeline: plasma::Kernel,
    uniform_buffer: Buffer,
    upload_scheme: Scheme,
    uniform_deposit: DepositTransaction,
    start_time: Instant,
    vsync: bool,
    frame_count: u32,
}

fn record_scheme(
    scheme: &mut Scheme,
    kernel: &plasma::Kernel,
    uniform: &Buffer,
    width: u32,
    height: u32,
    surface: Option<&SurfaceExchange>,
    readback: Option<&Texture>,
) -> Result<Option<Transaction>> {
    if let Some(surface) = surface {
        let (lease, present) = surface.bind_destination(scheme)?;
        kernel.record(scheme, "compute", uniform, &lease).over_2d(width, height);
        Ok(Some(present))
    } else {
        let target = readback.expect("capture readback");
        kernel.record(scheme, "compute", uniform, target).over_2d(width, height);

        Ok(None)
    }
}

fn output_size(state: &RenderState) -> (u32, u32) {
    if let Some(surface) = &state.surface {
        surface.size()
    } else {
        state.capture.as_ref().expect("capture").size()
    }
}

fn rebuild_scheme(state: &mut RenderState, width: u32, height: u32) {
    let mut scheme = Scheme::new(&state.ctx);
    let present = record_scheme(
        &mut scheme,
        &state.compute_pipeline,
        &state.uniform_buffer,
        width,
        height,
        state.surface.as_ref(),
        state.readback.as_ref(),
    )
    .expect("failed to record scheme");
    state.present = present;
    state.scheme = scheme;
}

impl App {
    fn init(&mut self, window: Option<Arc<Window>>) -> Result<()> {
        let warmup = self
            .warmup
            .take()
            .ok_or_else(|| anyhow::anyhow!("GPU warmup state missing"))?;
        let GpuWarmup {
            ctx,
            kernel: compute_pipeline,
            device,
            ..
        } = warmup;

        let (surface, capture, readback, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let (width, height) = surface.size();
            (Some(surface), None, None, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (None, Some(capture), Some(readback), width, height)
        };

        let uniform_buffer = device.acquire_buffer_with_data(
            &[Uniforms {
                width,
                height,
                time: 0.0,
            }],
            BufferKind::Scattered,
        )?;

        let mut scheme = Scheme::new(&ctx);
        let present = record_scheme(
            &mut scheme,
            &compute_pipeline,
            &uniform_buffer,
            width,
            height,
            surface.as_ref(),
            readback.as_ref(),
        )?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<Uniforms>(&uniform_buffer, 1),
        )?;

        self.state = Some(RenderState {
            window,
            ctx,
            surface,
            present,
            capture,
            readback,
            scheme,
            compute_pipeline,
            uniform_buffer,
            upload_scheme,
            uniform_deposit,
            start_time: Instant::now(),
            vsync: true,
            frame_count: 0,
        });

        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_some() {
            return;
        }
        let attrs = common::hidden_window("Goldy — Compute to Surface", INITIAL_WIDTH, INITIAL_HEIGHT);

        let window = Arc::new(event_loop.create_window(attrs).unwrap());

        if let Err(e) = self.init(Some(window.clone())) {
            tracing::error!("Failed to initialize: {}", e);
            event_loop.exit();
            return;
        }

        if let Some(state) = &mut self.state {
            if let Err(e) = render_frame(state) {
                tracing::error!("First frame error: {e}");
            }
        }
        common::reveal_window(&window);
        window.request_redraw();
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        let Some(state) = &mut self.state else {
            return;
        };

        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => match event.logical_key.as_ref() {
                Key::Named(NamedKey::Escape) => event_loop.exit(),
                Key::Character("v") => {
                    state.vsync = !state.vsync;
                    let mode = if state.vsync {
                        PresentMode::Fifo
                    } else {
                        PresentMode::Immediate
                    };
                    if let Some(surface) = &state.surface {
                        if let Err(e) = surface.set_present_mode(mode) {
                            eprintln!("Failed to set present mode: {e}");
                        } else {
                            println!(
                                "Vsync: {} (present mode: {:?})",
                                if state.vsync { "ON" } else { "OFF" },
                                mode
                            );
                        }
                    }
                }
                _ => {}
            },
            WindowEvent::Resized(new_size) if new_size.width > 0 && new_size.height > 0 => {
                if let Some(surface) = &state.surface {
                    let _ = surface.resize(new_size.width, new_size.height);
                    rebuild_scheme(state, new_size.width, new_size.height);
                }
                if let Some(window) = &state.window {
                    window.request_redraw();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = render_frame(state) {
                    tracing::error!("Render error: {}", e);
                }
                if let Some(window) = &state.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn render_frame(state: &mut RenderState) -> Result<()> {
    state.frame_count += 1;

    let (width, height) = output_size(state);
    if width == 0 || height == 0 {
        return Ok(());
    }

    let elapsed = state
        .capture
        .as_ref()
        .map(CaptureDump::time)
        .unwrap_or_else(|| state.start_time.elapsed().as_secs_f32());
    let uniforms = Uniforms {
        width,
        height,
        time: elapsed,
    };

    (&state.uniform_deposit << &uniforms)?;
    state.upload_scheme.submit()?;

    let mut submission = state.scheme.submit()?;
    if let Some(present) = &state.present {
        (&mut submission >> present).take()?;
    } else {
        let pixels = (&mut submission >> state.readback.as_ref().unwrap())
            .take::<u8>()?
            .to_vec();
        state.capture.as_mut().unwrap().write_rgba(&pixels)?;
    }

    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

The compute kernel is authored in Rust (#[goldy::compute]) in the example above.

ray_query

Builds a BLAS for one triangle plus a TLAS, then traces primary rays with inline RayQuery from a [goldy_compute] entry point and writes hits straight into the swapchain. No ray tracing pipeline or shader binding table is involved.

cargo run --features examples --example ray_query

What it demonstrates

  • Acceleration structure build (BLAS and TLAS) inside a scheme
  • Inline ray query from a compute entry point
  • Compute-to-surface output

Notes

The example exits 0 when RuntimeCapabilities::ray_query is false, and on the WebGPU backend, where Slang's WGSL target has no TraceRayInline.

Source

examples/ray_query.rs:

//! Compute ray query — a TLAS of one triangle, primary rays into the swapchain.
//!
//! Skips (exit 0) when `RuntimeCapabilities::ray_query` is false, or on WebGPU
//! (Slang WGSL has no `TraceRayInline`).
//!
//! Run with: cargo run --example ray_query --features examples

use anyhow::Result;
use goldy::{
    types::{BackendType, BufferFlags},
    AccelInstance, AccelerationStructure, Buffer, BufferKind, ComputePipeline, DepositTarget, DepositTransaction,
    Instance, MemoryExchange, NodeAccess, RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule,
    SurfaceConfig, SurfaceExchange, Texture, Transaction,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const RAY_SHADER: &str = r#"
import goldy_exp;


[goldy_compute]
[numthreads(8, 8, 1)]
void cs_main(BufRO<Uniforms> uniforms_buf, Accel scene, DirectSpatial<float4> output, ThreadId tid) {
    Uniforms u = uniforms_buf[0];
    if (tid.x >= u.width || tid.y >= u.height)
        return;

    float2 uv = (float2(tid.xy) + 0.5) / float2(u.width, u.height);
    float2 ndc = uv * 2.0 - 1.0;
    ndc.y = -ndc.y;

    RayDesc ray;
    ray.Origin = float3(0.0, 0.0, -2.0);
    ray.TMin = 0.001;
    ray.Direction = normalize(float3(ndc.x, ndc.y, 1.0));
    ray.TMax = 100.0;

    RayQuery<RAY_FLAG_FORCE_OPAQUE> q;
    q.TraceRayInline(scene, RAY_FLAG_FORCE_OPAQUE, 0xFF, ray);
    q.Proceed();

    float3 col = float3(0.05, 0.06, 0.12);
    if (q.CommittedStatus() == COMMITTED_TRIANGLE_HIT) {
        float2 bary = q.CommittedTriangleBarycentrics();
        col = float3(bary.x, bary.y, 1.0 - bary.x - bary.y);
        col += 0.15 * sin(u.time);
    }
    output[tid.xy] = float4(col, 1.0);
}
"#;

#[goldy::gpu]
struct Uniforms {
    width: u32,
    height: u32,
    time: f32,
}

const INITIAL_WIDTH: u32 = 800;
const INITIAL_HEIGHT: u32 = 600;

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let warmup = warm_gpu()?;
        let mut app = App {
            warmup: Some(warmup),
            state: None,
        };
        app.init(None)?;
        let state = app.state.as_mut().expect("capture state");
        while !state.capture.as_ref().is_none_or(CaptureDump::finished) {
            render_frame(state)?;
        }
        return Ok(());
    }

    println!("Goldy — Compute Ray Query");
    println!("=========================");
    println!("Press Escape to exit\n");

    let warmup = warm_gpu()?;
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    let mut app = App {
        warmup: Some(warmup),
        state: None,
    };
    event_loop.run_app(&mut app)?;
    Ok(())
}

struct GpuWarmup {
    ctx: goldy::Context,
    compute_pipeline: ComputePipeline,
    device: Arc<goldy::Runtime>,
    verts: Buffer,
    blas: AccelerationStructure,
    tlas: AccelerationStructure,
}

fn warm_gpu() -> Result<GpuWarmup> {
    let instance = Instance::new()?;
    let device = Arc::new(
        instance
            .request_adapter(&RequestAdapterOptions::default())?
            .request_runtime(&RuntimeDescriptor::default())?,
    );
    if !device.capabilities().ray_query {
        println!("skip: RuntimeCapabilities::ray_query is false on this adapter");
        std::process::exit(0);
    }
    if device.backend_type() == BackendType::WebGpu {
        println!("skip: WebGPU Slang path has no TraceRayInline");
        std::process::exit(0);
    }
    let ctx = device.create_context()?;
    let shader = ShaderModule::from_slang_with_gpu_types(&device, RAY_SHADER, &[Uniforms::GPU_TYPE])?;
    let compute_pipeline = ComputePipeline::new(&device, &shader)?;
    let positions: [[f32; 3]; 3] = [[0.0, 0.5, 0.0], [-0.7, -0.5, 0.0], [0.7, -0.5, 0.0]];
    let verts =
        device.acquire_buffer_with_data_and_flags(&positions, BufferKind::Scattered, BufferFlags::ACCEL_INPUT)?;
    let blas = AccelerationStructure::blas_triangles(&device, 1, 3, 12)?;
    let tlas = AccelerationStructure::tlas(&device, 1)?;
    Ok(GpuWarmup {
        ctx,
        compute_pipeline,
        device,
        verts,
        blas,
        tlas,
    })
}

#[derive(Default)]
struct App {
    warmup: Option<GpuWarmup>,
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scheme: Scheme,
    compute_pipeline: ComputePipeline,
    verts: Buffer,
    blas: AccelerationStructure,
    tlas: AccelerationStructure,
    uniform_buffer: Buffer,
    upload_scheme: Scheme,
    uniform_deposit: DepositTransaction,
    start_time: Instant,
    frame_count: u32,
}

#[allow(clippy::too_many_arguments)]
fn record_scheme(
    scheme: &mut Scheme,
    pipeline: &ComputePipeline,
    uniform: &Buffer,
    verts: &Buffer,
    blas: &AccelerationStructure,
    tlas: &AccelerationStructure,
    width: u32,
    height: u32,
    surface: Option<&SurfaceExchange>,
    readback: Option<&Texture>,
) -> Result<Option<Transaction>> {
    scheme.build_blas(blas, verts.whole(), 3, 12, None)?;
    let identity = [1.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0];
    scheme.build_tlas(
        tlas,
        &[AccelInstance {
            blas,
            transform: identity,
            mask: 0xFF,
            custom_index: 0,
        }],
    )?;
    let wg_x = width.div_ceil(8);
    let wg_y = height.div_ceil(8);
    if let Some(surface) = surface {
        let (lease, present) = surface.bind_destination(scheme)?;
        scheme
            .node("rays", pipeline)
            .with_parcel(uniform, NodeAccess::Read)
            .with_parcel(tlas, NodeAccess::Read)
            .with_present(&lease)
            .dispatch(wg_x, wg_y, 1);
        Ok(Some(present))
    } else {
        let target = readback.expect("capture readback");
        scheme
            .node("rays", pipeline)
            .with_parcel(uniform, NodeAccess::Read)
            .with_parcel(tlas, NodeAccess::Read)
            .with_parcel(target, NodeAccess::Write)
            .dispatch(wg_x, wg_y, 1);

        Ok(None)
    }
}

fn output_size(state: &RenderState) -> (u32, u32) {
    if let Some(surface) = &state.surface {
        surface.size()
    } else {
        state.capture.as_ref().expect("capture").size()
    }
}

fn rebuild_scheme(state: &mut RenderState, width: u32, height: u32) {
    let mut scheme = Scheme::new(&state.ctx);
    let present = record_scheme(
        &mut scheme,
        &state.compute_pipeline,
        &state.uniform_buffer,
        &state.verts,
        &state.blas,
        &state.tlas,
        width,
        height,
        state.surface.as_ref(),
        state.readback.as_ref(),
    )
    .expect("failed to record scheme");
    state.present = present;
    state.scheme = scheme;
}

impl App {
    fn init(&mut self, window: Option<Arc<Window>>) -> Result<()> {
        let warmup = self
            .warmup
            .take()
            .ok_or_else(|| anyhow::anyhow!("GPU warmup state missing"))?;
        let GpuWarmup {
            ctx,
            compute_pipeline,
            device,
            verts,
            blas,
            tlas,
        } = warmup;

        let (surface, capture, readback, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let (width, height) = surface.size();
            (Some(surface), None, None, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (None, Some(capture), Some(readback), width, height)
        };

        let uniform_buffer = device.acquire_buffer_with_data(
            &[Uniforms {
                width,
                height,
                time: 0.0,
            }],
            BufferKind::Scattered,
        )?;

        let mut scheme = Scheme::new(&ctx);
        let present = record_scheme(
            &mut scheme,
            &compute_pipeline,
            &uniform_buffer,
            &verts,
            &blas,
            &tlas,
            width,
            height,
            surface.as_ref(),
            readback.as_ref(),
        )?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer(&uniform_buffer, std::mem::size_of::<Uniforms>() as u64),
        )?;

        self.state = Some(RenderState {
            window,
            ctx,
            surface,
            present,
            capture,
            readback,
            scheme,
            compute_pipeline,
            verts,
            blas,
            tlas,
            uniform_buffer,
            upload_scheme,
            uniform_deposit,
            start_time: Instant::now(),
            frame_count: 0,
        });
        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_some() {
            return;
        }
        let attrs = common::hidden_window("Goldy — Compute Ray Query", INITIAL_WIDTH, INITIAL_HEIGHT);
        let window = Arc::new(event_loop.create_window(attrs).unwrap());
        if let Err(e) = self.init(Some(window.clone())) {
            tracing::error!("Failed to initialize: {}", e);
            event_loop.exit();
            return;
        }
        if let Some(state) = &mut self.state {
            if let Err(e) = render_frame(state) {
                tracing::error!("First frame error: {e}");
            }
        }
        common::reveal_window(&window);
        window.request_redraw();
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        let Some(state) = &mut self.state else {
            return;
        };
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key.as_ref(), Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::Resized(new_size) if new_size.width > 0 && new_size.height > 0 => {
                if let Some(surface) = &state.surface {
                    let _ = surface.resize(new_size.width, new_size.height);
                    rebuild_scheme(state, new_size.width, new_size.height);
                }
                if let Some(window) = &state.window {
                    window.request_redraw();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = render_frame(state) {
                    tracing::error!("Render error: {}", e);
                }
                if let Some(window) = &state.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn render_frame(state: &mut RenderState) -> Result<()> {
    state.frame_count += 1;
    let (width, height) = output_size(state);
    if width == 0 || height == 0 {
        return Ok(());
    }
    let uniforms = Uniforms {
        width,
        height,
        time: state
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| state.start_time.elapsed().as_secs_f32()),
    };
    (&state.uniform_deposit << &uniforms)?;
    state.upload_scheme.submit()?;
    let mut submission = state.scheme.submit()?;
    if let Some(present) = &state.present {
        (&mut submission >> present).take()?;
    } else {
        let pixels = (&mut submission >> state.readback.as_ref().unwrap())
            .take::<u8>()?
            .to_vec();
        state.capture.as_mut().unwrap().write_rgba(&pixels)?;
    }
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

The Slang source is inline in the example above.

tensor_algebra

Headless dense tensor algebra: fill a vector, then a tensor GEMV through the same retained scheme. There is no window and no capture clip.

cargo run --example tensor_algebra

CUDA compute-only (no graphics):

cargo run --no-default-features --features cuda,tensor --example tensor_algebra

What it demonstrates

  • TensorKernels / TensorRecorder recording into an ordinary Scheme
  • Checked views and semantic matmul (cuBLAS / MPS / stdlib)
  • Host observation through MemoryExchange

Source

examples/tensor_algebra.rs:

//! Headless dense tensor algebra: fill, add with broadcast, and GEMV.

use goldy::{
    Instance, RequestAdapterOptions, RuntimeDescriptor, Scheme, Tensor, TensorDType, TensorKernels, TensorScalar,
    TensorShape,
};
use std::ops::Shr;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let runtime = Instance::new()?
        .request_adapter(&RequestAdapterOptions::default())?
        .request_runtime(&RuntimeDescriptor::default())?;
    let ctx = runtime.create_context()?;
    let kernels = TensorKernels::new(&runtime)?;
    let x = Tensor::from_f32(&runtime, TensorShape::vector(4), &[1.0, 2.0, 3.0, 4.0])?;
    let w = Tensor::from_f32(
        &runtime,
        TensorShape::matrix(2, 4),
        &[1.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0],
    )?;
    let mut scheme = Scheme::new(&ctx);
    let y = {
        let mut rec = kernels.recorder(&mut scheme);
        rec.fill("fill_bias", x.view(), TensorScalar::F32(1.0))?;
        rec.matmul("gemv", w.view(), x.view())?
    };
    drop(kernels);

    let mut sub = scheme.submit()?;
    let bytes = (&mut sub >> y.buffer()).take::<u8>()?;
    let out: &[f32] = bytemuck::cast_slice(&bytes);
    println!("gemv(W, ones) = {out:?}");
    assert_eq!(out, &[1.0, 1.0]);
    let _ = TensorDType::F32;
    Ok(())
}

solid_cube

A solid 3D cube with per-face colours, indexed geometry, and a depth attachment on the scheme-leased render target.

cargo run --features examples --example solid_cube

What it demonstrates

  • Indexed draws
  • Depth attachment on a leased render target
  • CPU-side model/view/projection transforms uploaded per frame

Source

examples/solid_cube.rs:

//! Solid cube example - 3D filled cube with painter's algorithm.
//!
//! Demonstrates indexed rendering with 3D transformation via retained scheme.
//!
//! Run with: cargo run --example solid_cube

use goldy::{
    Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, IndexFormat, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc,
    RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad,
    Texture, TextureFormat, Transaction, Vertex2D,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[derive(Clone, Copy)]
struct Vertex3D {
    position: [f32; 3],
    color: Color,
}

fn generate_cube_vertices() -> Vec<Vertex3D> {
    let face_colors = [
        Color {
            r: 1.0,
            g: 0.3,
            b: 0.3,
            a: 1.0,
        },
        Color {
            r: 0.3,
            g: 1.0,
            b: 0.3,
            a: 1.0,
        },
        Color {
            r: 0.3,
            g: 0.3,
            b: 1.0,
            a: 1.0,
        },
        Color {
            r: 1.0,
            g: 1.0,
            b: 0.3,
            a: 1.0,
        },
        Color {
            r: 1.0,
            g: 0.3,
            b: 1.0,
            a: 1.0,
        },
        Color {
            r: 0.3,
            g: 1.0,
            b: 1.0,
            a: 1.0,
        },
    ];

    let faces: [[[f32; 3]; 4]; 6] = [
        [
            [-1.0, -1.0, -1.0],
            [1.0, -1.0, -1.0],
            [1.0, 1.0, -1.0],
            [-1.0, 1.0, -1.0],
        ],
        [[1.0, -1.0, 1.0], [-1.0, -1.0, 1.0], [-1.0, 1.0, 1.0], [1.0, 1.0, 1.0]],
        [
            [-1.0, -1.0, 1.0],
            [-1.0, -1.0, -1.0],
            [-1.0, 1.0, -1.0],
            [-1.0, 1.0, 1.0],
        ],
        [[1.0, -1.0, -1.0], [1.0, -1.0, 1.0], [1.0, 1.0, 1.0], [1.0, 1.0, -1.0]],
        [[-1.0, 1.0, -1.0], [1.0, 1.0, -1.0], [1.0, 1.0, 1.0], [-1.0, 1.0, 1.0]],
        [
            [-1.0, -1.0, 1.0],
            [1.0, -1.0, 1.0],
            [1.0, -1.0, -1.0],
            [-1.0, -1.0, -1.0],
        ],
    ];

    let mut vertices = Vec::new();
    for (face_idx, face) in faces.iter().enumerate() {
        for &pos in face {
            vertices.push(Vertex3D {
                position: pos,
                color: face_colors[face_idx],
            });
        }
    }
    vertices
}

fn rotate_y(p: [f32; 3], angle: f32) -> [f32; 3] {
    let (s, c) = (angle.sin(), angle.cos());
    [p[0] * c + p[2] * s, p[1], -p[0] * s + p[2] * c]
}

fn rotate_x(p: [f32; 3], angle: f32) -> [f32; 3] {
    let (s, c) = (angle.sin(), angle.cos());
    [p[0], p[1] * c - p[2] * s, p[1] * s + p[2] * c]
}

fn project(p: [f32; 3], fov: f32) -> [f32; 2] {
    let z = p[2] + 4.0;
    let scale = fov / z;
    [p[0] * scale, p[1] * scale]
}

const MAX_CUBE_VERTICES: usize = 24;
const MAX_CUBE_INDICES: usize = 36;

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    vertex_parcel: Option<Buffer>,
    index_parcel: Option<Buffer>,
    upload_scheme: Option<Scheme>,
    vertex_deposit: Option<DepositTransaction>,
    index_deposit: Option<DepositTransaction>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    start_time: Instant,
    cube_vertices: Vec<Vertex3D>,
    frame_count: u64,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            start_time: Instant::now(),
            vertex_parcel: None,
            index_parcel: None,
            upload_scheme: None,
            vertex_deposit: None,
            index_deposit: None,
            cube_vertices: generate_cube_vertices(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: Vertex2D::layout(),
                topology: PrimitiveTopology::TriangleList,
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        vertex_parcel: &Buffer,
        index_parcel: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass(
            "solid_cube",
            scene_rt,
            TargetLoad::Clear(Color {
                r: 0.02,
                g: 0.02,
                b: 0.05,
                a: 1.0,
            }),
        );
        pass.with_parcel(vertex_parcel, NodeAccess::Read);
        pass.with_parcel(index_parcel, NodeAccess::Read);

        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, vertex_parcel);
        pass.set_index_buffer(index_parcel, IndexFormat::Uint16);
        pass.draw_indexed(0..MAX_CUBE_INDICES as u32, 0, 0..1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, goldy::shader::builtins::VERTEX_COLOR_2D)?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let vertex_parcel = device.acquire_buffer_sized::<Vertex2D>(
            MAX_CUBE_VERTICES as u64,
            BufferKind::Scattered,
            BufferFlags::empty(),
        )?;
        let index_parcel =
            device.acquire_buffer_sized::<u16>(MAX_CUBE_INDICES as u64, BufferKind::Scattered, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &vertex_parcel, &index_parcel, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        let ctx = self.ctx.as_ref().unwrap();
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.vertex_parcel = Some(vertex_parcel);
        self.index_parcel = Some(index_parcel);
        let vertex_parcel = self.vertex_parcel.as_ref().unwrap();
        let index_parcel = self.index_parcel.as_ref().unwrap();
        let mut upload_scheme = Scheme::new(ctx);
        let memory = MemoryExchange::new(ctx);
        let vertex_deposit = memory.bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer(vertex_parcel, vertex_parcel.byte_size()),
        )?;
        let index_deposit = memory.bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer(index_parcel, index_parcel.byte_size()),
        )?;
        self.upload_scheme = Some(upload_scheme);
        self.vertex_deposit = Some(vertex_deposit);
        self.index_deposit = Some(index_deposit);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());

        let rotated_3d: Vec<[f32; 3]> = self
            .cube_vertices
            .iter()
            .map(|v| rotate_x(rotate_y(v.position, time), time * 0.7))
            .collect();

        let mut face_depths: Vec<(usize, f32)> = (0..6)
            .map(|face_idx| {
                let base = face_idx * 4;
                let avg_z =
                    (rotated_3d[base][2] + rotated_3d[base + 1][2] + rotated_3d[base + 2][2] + rotated_3d[base + 3][2])
                        / 4.0;
                (face_idx, avg_z)
            })
            .collect();
        face_depths.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap());

        let mut vertices = Vec::with_capacity(24);
        let mut sorted_indices = Vec::with_capacity(36);

        for (new_base, (face_idx, _)) in face_depths.iter().enumerate() {
            let old_base = face_idx * 4;
            let new_base = (new_base * 4) as u16;

            for i in 0..4 {
                let projected = project(rotated_3d[old_base + i], 2.0);
                vertices.push(Vertex2D::new(
                    projected[0],
                    projected[1],
                    self.cube_vertices[old_base + i].color,
                ));
            }

            sorted_indices.extend_from_slice(&[
                new_base,
                new_base + 1,
                new_base + 2,
                new_base,
                new_base + 2,
                new_base + 3,
            ]);
        }

        let upload = self.upload_scheme.as_mut().unwrap();
        (self.vertex_deposit.as_ref().unwrap() << vertices.as_slice())?;
        (self.index_deposit.as_ref().unwrap() << sorted_indices.as_slice())?;
        upload.submit()?;

        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        self.frame_count += 1;
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(vb), Some(ib)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.vertex_parcel.as_ref(),
            self.index_parcel.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, vb, ib, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Solid Cube (Scheme + Present)", 800, 800))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Solid Cube Example (Scheme + Present) - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/vertex_color_2d.slang:

// Simple 2D vertex + fragment shader for colored vertices.
// Used by: triangle, particles, starfield, bouncing_lines, spinning_cube, instancing, waveform

struct VertexInput {
    float2 position : POSITION;
    float4 color : COLOR;
};

struct VertexOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[goldy_vertex]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.color = input.color;
    return output;
}

[goldy_fragment]
float4 fs_main(VertexOutput input) : SV_Target {
    return input.color;
}

spinning_cube

A wireframe cube drawn with LINE_LIST topology and a hand-rolled 3D projection, which keeps the example free of any matrix library.

cargo run --features examples --example spinning_cube

What it demonstrates

  • Line primitives in a render pipeline
  • Per-frame vertex uploads through a deposit transaction

Source

examples/spinning_cube.rs:

//! Spinning cube example - 3D wireframe cube.
//!
//! Demonstrates 3D projection via retained scheme with copy-to-present.
//!
//! Run with: cargo run --example spinning_cube

use goldy::{
    Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease, LeaseRenderTarget,
    MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, Vertex2D,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const CUBE_VERTICES: [[f32; 3]; 8] = [
    [-1.0, -1.0, -1.0],
    [1.0, -1.0, -1.0],
    [1.0, 1.0, -1.0],
    [-1.0, 1.0, -1.0],
    [-1.0, -1.0, 1.0],
    [1.0, -1.0, 1.0],
    [1.0, 1.0, 1.0],
    [-1.0, 1.0, 1.0],
];

const CUBE_EDGES: [[usize; 2]; 12] = [
    [0, 1],
    [1, 2],
    [2, 3],
    [3, 0],
    [4, 5],
    [5, 6],
    [6, 7],
    [7, 4],
    [0, 4],
    [1, 5],
    [2, 6],
    [3, 7],
];

fn rotate_y(p: [f32; 3], angle: f32) -> [f32; 3] {
    let (s, c) = (angle.sin(), angle.cos());
    [p[0] * c + p[2] * s, p[1], -p[0] * s + p[2] * c]
}

fn rotate_x(p: [f32; 3], angle: f32) -> [f32; 3] {
    let (s, c) = (angle.sin(), angle.cos());
    [p[0], p[1] * c - p[2] * s, p[1] * s + p[2] * c]
}

fn project(p: [f32; 3], fov: f32) -> [f32; 2] {
    let z = p[2] + 4.0;
    let scale = fov / z;
    [p[0] * scale, p[1] * scale]
}

const MAX_LINE_VERTICES: usize = CUBE_EDGES.len() * 2;

struct App {
    instance: Instance,
    // Window + surface before ctx/device so Escape teardown destroys the swapchain
    // while the CUDA context stream is still alive (avoids cudarc sticky error_state
    // from stream Drop racing destroy_surface bind).
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    vertex_deposit: Option<DepositTransaction>,
    vertex_parcel: Option<Buffer>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            vertex_deposit: None,
            vertex_parcel: None,
            pipeline: None,
            shader: None,
            ctx: None,
            device: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: Vertex2D::layout(),
                topology: PrimitiveTopology::LineList,
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        vertex_parcel: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass(
            "spinning_cube",
            scene_rt,
            TargetLoad::Clear(Color {
                r: 0.02,
                g: 0.02,
                b: 0.05,
                a: 1.0,
            }),
        );
        pass.with_parcel(vertex_parcel, NodeAccess::Read);

        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, vertex_parcel);
        pass.draw(0..MAX_LINE_VERTICES as u32, 0..1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, goldy::shader::builtins::VERTEX_COLOR_2D)?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let vertex_parcel = device.acquire_buffer_sized::<Vertex2D>(
            MAX_LINE_VERTICES as u64,
            BufferKind::Scattered,
            BufferFlags::empty(),
        )?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &vertex_parcel, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        let ctx = self.ctx.as_ref().unwrap();
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.vertex_parcel = Some(vertex_parcel);
        let vertex_parcel = self.vertex_parcel.as_ref().unwrap();
        let mut upload_scheme = Scheme::new(ctx);
        let vertex_deposit = MemoryExchange::new(ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer(vertex_parcel, vertex_parcel.byte_size()),
        )?;
        self.upload_scheme = Some(upload_scheme);
        self.vertex_deposit = Some(vertex_deposit);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());
        let transformed: Vec<[f32; 3]> = CUBE_VERTICES
            .iter()
            .map(|&v| rotate_x(rotate_y(v, time), time * 0.7))
            .collect();

        let mut vertices: Vec<Vertex2D> = Vec::new();
        for edge in &CUBE_EDGES {
            let p1 = project(transformed[edge[0]], 2.0);
            let p2 = project(transformed[edge[1]], 2.0);

            let z1 = transformed[edge[0]][2];
            let z2 = transformed[edge[1]][2];
            let avg_z = (z1 + z2) / 2.0;
            let brightness = (avg_z + 1.5) / 3.0;
            let color = Color {
                r: 0.2 + brightness * 0.8,
                g: 0.5 + brightness * 0.5,
                b: 1.0,
                a: 1.0,
            };

            vertices.push(Vertex2D::new(p1[0], p1[1], color));
            vertices.push(Vertex2D::new(p2[0], p2[1], color));
        }

        let upload = self.upload_scheme.as_mut().unwrap();
        (self.vertex_deposit.as_ref().unwrap() << vertices.as_slice())?;
        upload.submit()?;

        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(vertex_parcel)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.vertex_parcel.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, vertex_parcel, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Spinning Cube (Scheme + Present)",
                        800,
                        800,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Spinning Cube Example (Scheme + Present) - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/vertex_color_2d.slang:

// Simple 2D vertex + fragment shader for colored vertices.
// Used by: triangle, particles, starfield, bouncing_lines, spinning_cube, instancing, waveform

struct VertexInput {
    float2 position : POSITION;
    float4 color : COLOR;
};

struct VertexOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[goldy_vertex]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.color = input.color;
    return output;
}

[goldy_fragment]
float4 fs_main(VertexOutput input) : SV_Target {
    return input.color;
}

depth_quads

Two fullscreen quads whose depths cross periodically. Because depth testing decides visibility, the picture is independent of the order the quads are drawn in — which is exactly what the animation makes visible.

cargo run --features examples --example depth_quads

What it demonstrates

  • Context::lease_render_target with a depth attachment
  • Depth-stencil state on a render pipeline
  • Draw-order independence

Source

examples/depth_quads.rs:

//! Depth quads example - two fullscreen quads whose depths cross periodically.
//!
//! Depth-tested rendering via an offscreen context-leased render target with a depth attachment
//! (`Context::lease_render_target` with depth), then copy-to-present through a retained scheme.
//!
//! Run with: cargo run --example depth_quads

use goldy::{
    Buffer, BufferFlags, BufferKind, Color, CompareFunction, DepositTarget, DepositTransaction, DepthFormat,
    DepthStencilState, Instance, Lease, LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline,
    RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange,
    TargetLoad, Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct DepthVertex {
    position: [f32; 3],
    color: [f32; 4],
}

impl DepthVertex {
    const fn new(x: f32, y: f32, z: f32, r: f32, g: f32, b: f32) -> Self {
        Self {
            position: [x, y, z],
            color: [r, g, b, 1.0],
        }
    }
}

fn depth_vertex_layout() -> VertexBufferLayout {
    DepthVertex::GPU_TYPE
        .vertex_buffer_layout()
        .expect("depth vertex layout")
}

#[allow(clippy::too_many_arguments)]
fn quad_verts(x0: f32, y0: f32, x1: f32, y1: f32, z: f32, r: f32, g: f32, b: f32) -> [DepthVertex; 6] {
    let tl = DepthVertex::new(x0, y1, z, r, g, b);
    let bl = DepthVertex::new(x0, y0, z, r, g, b);
    let br = DepthVertex::new(x1, y0, z, r, g, b);
    let tr = DepthVertex::new(x1, y1, z, r, g, b);
    [tl, bl, br, tl, br, tr]
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    warm_parcel: Option<Buffer>,
    cool_parcel: Option<Buffer>,
    warm_deposit: Option<DepositTransaction>,
    cool_deposit: Option<DepositTransaction>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    window: Option<Arc<Window>>,
    frame_count: u64,
    start_time: std::time::Instant,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            warm_parcel: None,
            cool_parcel: None,
            warm_deposit: None,
            cool_deposit: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            window: None,
            frame_count: 0,
            start_time: std::time::Instant::now(),
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: depth_vertex_layout(),
                depth_stencil: Some(DepthStencilState {
                    format: DepthFormat::Depth32Float,
                    depth_write_enabled: true,
                    depth_compare: CompareFunction::Less,
                }),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        warm_parcel: &Buffer,
        cool_parcel: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("depth_quads", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(warm_parcel, NodeAccess::Read);
        pass.with_parcel(cool_parcel, NodeAccess::Read);
        pass.clear_depth(1.0);
        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, warm_parcel);
        pass.draw(0..6, 0..1);
        pass.set_vertex_buffer(0, cool_parcel);
        pass.draw(0..6, 0..1);
        pass.finish();
    }

    fn bind_uploads(
        ctx: &goldy::Context,
        scheme: &mut Scheme,
        warm_parcel: &Buffer,
        cool_parcel: &Buffer,
    ) -> anyhow::Result<(DepositTransaction, DepositTransaction)> {
        let memory = MemoryExchange::new(ctx);
        let warm_deposit = memory.bind_deposit(scheme, DepositTarget::buffer(warm_parcel, warm_parcel.byte_size()))?;
        let cool_deposit = memory.bind_deposit(scheme, DepositTarget::buffer(cool_parcel, cool_parcel.byte_size()))?;
        Ok((warm_deposit, cool_deposit))
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, include_str!("../shaders/depth_test.slang"))?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let warm_parcel = device.acquire_buffer_sized::<DepthVertex>(6, BufferKind::Scattered, BufferFlags::empty())?;
        let cool_parcel = device.acquire_buffer_sized::<DepthVertex>(6, BufferKind::Scattered, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, Some(DepthFormat::Depth32Float))?;
        let (warm_deposit, cool_deposit) = Self::bind_uploads(&ctx, &mut scheme, &warm_parcel, &cool_parcel)?;
        Self::record_pass(&mut scheme, &pipeline, &warm_parcel, &cool_parcel, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.warm_parcel = Some(warm_parcel);
        self.cool_parcel = Some(cool_parcel);
        self.warm_deposit = Some(warm_deposit);
        self.cool_deposit = Some(cool_deposit);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let t = self.frame_count as f32 * 0.04;
        let warm_z = t.sin() * 0.4 + 0.5;
        let cool_z = (t * 1.3 + 1.0).sin() * 0.4 + 0.5;

        let warm_verts = quad_verts(-1.0, -1.0, 1.0, 1.0, warm_z, 0.95, 0.35, 0.1);
        let cool_verts = quad_verts(-1.0, -1.0, 1.0, 1.0, cool_z, 0.1, 0.6, 0.95);

        let winner = if warm_z < cool_z { "WARM wins" } else { "COOL wins" };
        if let Some(window) = &self.window {
            window.set_title(&format!(
                "Depth Quads  |  warm z={:.3}  cool z={:.3}  →  {}",
                warm_z, cool_z, winner
            ));
        }

        (self.warm_deposit.as_ref().unwrap() << warm_verts.as_slice())?;
        (self.cool_deposit.as_ref().unwrap() << cool_verts.as_slice())?;

        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        self.frame_count += 1;
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(warm), Some(cool)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.warm_parcel.as_ref(),
            self.cool_parcel.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) =
                        ctx.lease_render_target(width.max(1), height.max(1), format, Some(DepthFormat::Depth32Float))
                    {
                        if let Ok((warm_deposit, cool_deposit)) = Self::bind_uploads(ctx, &mut scheme, warm, cool) {
                            Self::record_pass(&mut scheme, pipeline, warm, cool, &rt);
                            if let Ok(present) =
                                Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                            {
                                self.warm_deposit = Some(warm_deposit);
                                self.cool_deposit = Some(cool_deposit);
                                self.present = present;
                                self.scheme = Some(scheme);
                                self.scene_rt = Some(rt);
                            }
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Depth Quads (Scheme + Present)",
                        900,
                        600,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Depth Quads Example (Scheme + Present)");
    println!("Press Escape or close window to exit.\n");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/depth_test.slang:

// 3D depth-testing shader: vertex carries (x, y, z) position and RGBA color.
// Used by depth occlusion screenshot tests.

struct VIn {
    float3 position : POSITION;
    float4 color : COLOR;
};

struct VOut {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[goldy_vertex]
VOut vs_main(VIn v) {
    VOut o;
    o.position = float4(v.position.xy, v.position.z, 1.0);
    o.color = v.color;
    return o;
}

[goldy_fragment]
float4 fs_main(VOut i) : SV_Target {
    return i.color;
}

textured_quad

A procedurally generated checkerboard texture sampled onto a quad. The vertex stage declares no bindless resources at all; only the fragment stage takes tex and smp, and Goldy binds them by pipeline name.

cargo run --features examples --example textured_quad

What it demonstrates

  • Stage-local resource declarations
  • Named draw bindings (tex, smp)
  • Automatic payload linking of FullscreenVarying between stages
  • Texture upload through a deposit transaction

Source

examples/textured_quad.rs:

//! Textured quad example — stage-local resources, automatic payload linking, named draw bindings.
//!
//! The vertex stage has no bindless resources; the fragment stage declares `tex` and `smp`.
//! Goldy links `FullscreenVarying` and binds resources by pipeline name.
//!
//! Run with: cargo run --example textured_quad

use goldy::{
    types::{AddressMode, FilterMode, SamplerDesc, TextureFlags, TextureFormat, TextureKind},
    Buffer, BufferKind, Color, Instance, Lease, LeaseRenderTarget, MemoryExchange, NodeAccess, Parcel, RenderPipeline,
    RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Sampler, Scheme, ShaderBinding, ShaderModule,
    SurfaceConfig, SurfaceExchange, TargetLoad, Texture, Transaction, Vertex2DUv,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const TEXTURED_SHADER: &str = r#"
import goldy_exp;

[goldy_vertex]
FullscreenVarying vs_main(FullscreenVertex input) {
    return vs_fullscreen(input);
}

[goldy_fragment]
float4 fs_main(Interpolated<float4> tex, Filter smp, FullscreenVarying input) : SV_Target {
    return tex.Sample(smp, input.uv);
}
"#;

/// Generate a checkerboard texture in RGBA8 format
fn generate_checkerboard(width: u32, height: u32, checker_size: u32) -> Vec<u8> {
    let mut data = Vec::with_capacity((width * height * 4) as usize);

    for y in 0..height {
        for x in 0..width {
            let checker_x = (x / checker_size) % 2;
            let checker_y = (y / checker_size) % 2;
            let is_white = (checker_x + checker_y).is_multiple_of(2);

            if is_white {
                data.extend_from_slice(&[255, 255, 255, 255]);
            } else {
                data.extend_from_slice(&[50, 100, 200, 255]);
            }
        }
    }

    data
}

const QUAD_VERTICES: [Vertex2DUv; 6] = [
    Vertex2DUv {
        position: [-1.0, -1.0],
        uv: [0.0, 1.0],
    },
    Vertex2DUv {
        position: [1.0, -1.0],
        uv: [1.0, 1.0],
    },
    Vertex2DUv {
        position: [1.0, 1.0],
        uv: [1.0, 0.0],
    },
    Vertex2DUv {
        position: [-1.0, -1.0],
        uv: [0.0, 1.0],
    },
    Vertex2DUv {
        position: [1.0, 1.0],
        uv: [1.0, 0.0],
    },
    Vertex2DUv {
        position: [-1.0, 1.0],
        uv: [0.0, 0.0],
    },
];

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    vertex_buffer: Option<Buffer>,
    texture: Option<Texture>,
    sampler: Option<Sampler>,
    start_time: Instant,
    frame_count: u64,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            vertex_buffer: None,
            texture: None,
            sampler: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: Vertex2DUv::layout(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        vertex_buffer: &Buffer,
        texture: &Parcel,
        sampler: &Sampler,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let bindings = [
            ShaderBinding::read("tex", texture),
            ShaderBinding::sampler("smp", sampler),
        ];

        let mut pass = scheme.render_pass(
            "textured_quad",
            scene_rt,
            TargetLoad::Clear(Color {
                r: 0.1,
                g: 0.1,
                b: 0.15,
                a: 1.0,
            }),
        );
        pass.with_shader_bindings(&bindings);
        pass.with_parcel(vertex_buffer, NodeAccess::Read);

        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, vertex_buffer);
        pass.draw(0..6, 0..1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, TEXTURED_SHADER)?;

        let tex_width = 256u32;
        let tex_height = 256u32;
        let checker_data = generate_checkerboard(tex_width, tex_height, 32);

        let texture = device.acquire_texture(
            tex_width,
            tex_height,
            TextureFormat::Rgba8Unorm,
            TextureKind::Interpolated,
            TextureFlags::COPY_DST,
            Some(&checker_data),
        )?;

        let sampler = Sampler::new(
            &device,
            &SamplerDesc {
                mag_filter: FilterMode::Linear,
                min_filter: FilterMode::Linear,
                mipmap_filter: FilterMode::Nearest,
                address_mode_u: AddressMode::Repeat,
                address_mode_v: AddressMode::Repeat,
                address_mode_w: AddressMode::Repeat,
                max_anisotropy: 1.0,
                compare: None,
                lod_min_clamp: 0.0,
                lod_max_clamp: 32.0,
            },
        )?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let vertex_buffer = device.acquire_buffer_with_data(&QUAD_VERTICES, BufferKind::Scattered)?;
        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &vertex_buffer, &texture, &sampler, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.vertex_buffer = Some(vertex_buffer);
        self.texture = Some(texture);
        self.sampler = Some(sampler);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        self.frame_count += 1;
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(vertex_buffer), Some(texture), Some(sampler)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.vertex_buffer.as_ref(),
            self.texture.as_ref(),
            self.sampler.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, vertex_buffer, texture, sampler, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Textured Quad (Scheme + Present)",
                        800,
                        800,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Textured Quad Example (Scheme + Present)");
    println!("Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

The Slang source is inline in the example above.

instancing

GPU-driven instancing: a compute pass updates per-instance transforms in a storage buffer, and one instanced draw renders them all. The per-instance layout lives in examples/instance2d.rs and mirrors QuadInstance in the Slang shaders.

cargo run --features examples --example instancing

What it demonstrates

  • Compute-updated instance data consumed by a raster pass
  • Matching host and shader struct layouts
  • One draw call for many objects

Source

examples/instancing.rs:

//! Instancing example - render many objects efficiently.
//!
//! Demonstrates retained scheme with compute dispatch → offscreen render → copy-to-present.
//!
//! Run with: cargo run --example instancing

use anyhow::Result;
use goldy::{
    Buffer, BufferFlags, BufferKind, Color, ComputePipeline, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc,
    RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad,
    Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;

mod instance2d;
use instance2d::Instance2D;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const GRID_SIZE: u32 = 20;
const QUAD_SIZE: f32 = 0.03;
const NUM_QUADS: u32 = GRID_SIZE * GRID_SIZE;

#[goldy::gpu]
struct AnimParams {
    time: f32,
    delta_time: f32,
    total_instances: u32,
}

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut state = RenderState::new(None)?;
        while !state.capture_done() {
            state.render()?;
        }
        return Ok(());
    }

    println!("Goldy Instancing Example - {} quads (Scheme + Present)", NUM_QUADS);
    println!("Press Escape to exit");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App::default();
    event_loop.run_app(&mut app)?;

    Ok(())
}

#[derive(Default)]
struct App {
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    device: Arc<goldy::Runtime>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    present: Option<Transaction>,
    scheme: Scheme,
    scene_rt: Lease<LeaseRenderTarget>,
    compute_pipeline: ComputePipeline,
    render_shader: ShaderModule,
    render_pipeline: RenderPipeline,
    instance_buffer: Buffer,
    params_buffer: Buffer,
    upload_scheme: Scheme,
    params_deposit: DepositTransaction,
    start_time: Instant,
    last_time: f32,
    frame_count: u32,
}

impl RenderState {
    fn create_render_pipeline(
        device: &goldy::Runtime,
        render_shader: &ShaderModule,
        format: TextureFormat,
    ) -> Result<RenderPipeline> {
        common::render_pipeline(
            device,
            render_shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                topology: PrimitiveTopology::TriangleList,
                ..Default::default()
            },
        )
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn record_scheme(
        scheme: &mut Scheme,
        compute_pipeline: &ComputePipeline,
        render_pipeline: &RenderPipeline,
        instance_buffer: &Buffer,
        params_buffer: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        scheme
            .node("update_instances", compute_pipeline)
            .with_parcel(instance_buffer, NodeAccess::ReadWrite)
            .with_parcel(params_buffer, NodeAccess::Read)
            .dispatch(NUM_QUADS.div_ceil(64), 1, 1);

        let bg_color = Color {
            r: 0.02,
            g: 0.02,
            b: 0.04,
            a: 1.0,
        };

        let mut pass = scheme.render_pass("instancing", scene_rt, TargetLoad::Clear(bg_color));
        pass.with_parcel(instance_buffer, NodeAccess::Read);
        pass.set_pipeline(render_pipeline);
        pass.draw(0..6, 0..NUM_QUADS);
        pass.finish();
    }

    fn target(&self) -> (TextureFormat, u32, u32) {
        if let Some(surface) = &self.surface {
            let (width, height) = surface.size();
            (surface.format(), width, height)
        } else {
            let capture = self.capture.as_ref().expect("capture dump");
            let (width, height) = capture.size();
            (CaptureDump::format(), width, height)
        }
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn rerecord_scheme(&mut self) {
        let mut scheme = Scheme::new(&self.ctx);
        let (format, width, height) = self.target();
        if let Ok(rt) = self.ctx.lease_render_target(width.max(1), height.max(1), format, None) {
            self.scene_rt = rt;
            Self::record_scheme(
                &mut scheme,
                &self.compute_pipeline,
                &self.render_pipeline,
                &self.instance_buffer,
                &self.params_buffer,
                &self.scene_rt,
            );
            if let Ok(present) = Self::bind_frame(
                &mut scheme,
                &self.scene_rt,
                self.surface.as_ref(),
                self.readback.as_ref(),
            ) {
                self.present = present;
                self.scheme = scheme;
            }
        }
    }

    fn new(window: Option<Arc<Window>>) -> Result<Self> {
        let instance = Instance::new()?;
        let device = Arc::new(
            instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let compute_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/instancing_update.slang"),
            &[Instance2D::GPU_TYPE, AnimParams::GPU_TYPE],
        )?;
        let render_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/instancing_render.slang"),
            &[Instance2D::GPU_TYPE],
        )?;

        let mut instances = Vec::with_capacity(NUM_QUADS as usize);
        for i in 0..GRID_SIZE {
            for j in 0..GRID_SIZE {
                let nx = (i as f32 / (GRID_SIZE - 1) as f32) * 2.0 - 1.0;
                let ny = (j as f32 / (GRID_SIZE - 1) as f32) * 2.0 - 1.0;
                instances.push(Instance2D::new(
                    nx * 0.85,
                    ny * 0.85,
                    0.0,
                    QUAD_SIZE,
                    [1.0, 1.0, 1.0, 1.0],
                ));
            }
        }

        let instance_buffer = device.acquire_buffer_with_data(&instances, BufferKind::Scattered)?;
        let params_buffer =
            device.acquire_buffer_sized::<AnimParams>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let compute_pipeline = ComputePipeline::new(&device, &compute_shader)?;
        let render_pipeline = Self::create_render_pipeline(&device, &render_shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_scheme(
            &mut scheme,
            &compute_pipeline,
            &render_pipeline,
            &instance_buffer,
            &params_buffer,
            &scene_rt,
        );
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let params_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<AnimParams>(&params_buffer, 1),
        )?;

        println!(
            "Created instancing example with {} quads (GPU compute + graphics)",
            NUM_QUADS
        );

        Ok(Self {
            window,
            device,
            ctx,
            surface,
            capture,
            readback,
            present,
            scheme,
            scene_rt,
            compute_pipeline,
            render_shader,
            render_pipeline,
            instance_buffer,
            params_buffer,
            upload_scheme,
            params_deposit,
            start_time: Instant::now(),
            last_time: 0.0,
            frame_count: 0,
        })
    }

    fn render(&mut self) -> Result<()> {
        self.frame_count += 1;

        let time = if let Some(capture) = &self.capture {
            capture.time()
        } else {
            self.start_time.elapsed().as_secs_f32()
        };
        let delta_time = time - self.last_time;
        self.last_time = time;

        let params = AnimParams {
            time,
            delta_time,
            total_instances: NUM_QUADS,
        };

        (&self.params_deposit << &params)?;
        self.upload_scheme.submit()?;

        let mut submission = self.scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        if let Some(window) = &self.window {
            window.request_redraw();
        }
        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        format!("Goldy - Instancing ({} quads, Scheme + Present)", NUM_QUADS),
                        800,
                        800,
                    ))
                    .expect("Failed to create window"),
            );

            match RenderState::new(Some(window.clone())) {
                Ok(mut state) => {
                    if let Err(e) = state.render() {
                        tracing::error!("First frame error: {e}");
                    }
                    common::reveal_window(&window);
                    self.state = Some(state);
                    window.request_redraw();
                }
                Err(e) => {
                    tracing::error!("Failed to create render state: {}", e);
                    event_loop.exit();
                }
            }
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::Resized(size) => {
                if let Some(state) = &mut self.state {
                    if size.width > 0 && size.height > 0 {
                        let Some(surface) = state.surface.as_ref() else {
                            return;
                        };
                        let (prev_w, prev_h) = surface.size();
                        if size.width == prev_w && size.height == prev_h {
                            return;
                        }
                        let _ = surface.resize(size.width, size.height);
                        let format = surface.format();
                        if let Ok(pipeline) =
                            RenderState::create_render_pipeline(&state.device, &state.render_shader, format)
                        {
                            state.render_pipeline = pipeline;
                        }
                        state.rerecord_scheme();
                    }
                }
            }
            WindowEvent::RedrawRequested => {
                if let Some(state) = &mut self.state {
                    if let Err(e) = state.render() {
                        tracing::error!("Render error: {}", e);
                    }
                }
            }
            _ => {}
        }
    }
}

The example pulls in examples/common.rs, examples/instance2d.rs — see Shared Helpers.

Shaders

shaders/instancing_update.slang:

// Instancing compute shader - updates rotation and color for each quad instance

import goldy_exp;


float3 hsv_to_rgb(float h, float s, float v) {
    float4 K = float4(1.0, 2.0/3.0, 1.0/3.0, 3.0);
    float3 p = abs(frac(h.xxx + K.xyz) * 6.0 - K.www);
    return v * lerp(K.xxx, clamp(p - K.xxx, 0.0, 1.0), s);
}

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Instance2D> INSTANCES, AnimParams params, ThreadId id) {
    uint idx = id.x;
    if (idx >= params.total_instances) return;
    
    Instance2D inst = INSTANCES[idx];
    
    float phase = (float(idx) / float(params.total_instances)) * 6.28318;
    inst.rotation = params.time * 2.0 + phase;
    
    float hue = frac(float(idx) / float(params.total_instances) + params.time * 0.1);
    inst.color = float4(hsv_to_rgb(hue, 0.8, 0.9), 1.0);
    
    INSTANCES[idx] = inst;
}

shaders/instancing_render.slang:

// Instancing render shader - renders quads from instance buffer

import goldy_exp;


struct VSOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[goldy_vertex]
VSOutput vs_main(Scattered<Instance2D> INSTANCES, VertexId vertex_id, InstanceId instance_id) {
    VSOutput output;
    
    Instance2D inst = INSTANCES[instance_id.value];
    float2 world_pos = quad_position_rotated(vertex_id.value, inst.position, inst.scale, inst.rotation);
    
    output.position = float4(world_pos, 0.0, 1.0);
    output.color = inst.color;
    
    return output;
}

[goldy_fragment]
float4 fs_main(VSOutput input) : SV_Target {
    return input.color;
}

bouncing_lines

Line segments bouncing off the window edges. A compute pass integrates the simple physics; a LINE_LIST pipeline draws the result.

cargo run --features examples --example bouncing_lines

What it demonstrates

  • LINE_LIST topology
  • Compute dispatch feeding a raster pass in one retained scheme

Source

examples/bouncing_lines.rs:

//! Bouncing lines example - animated lines bouncing off walls.
//!
//! Demonstrates retained scheme with compute dispatch → offscreen render → copy-to-present.
//!
//! Run with: cargo run --example bouncing_lines

use anyhow::Result;
use goldy::{
    Buffer, BufferKind, Color, ComputePipeline, Instance, Lease, LeaseRenderTarget, MemoryExchange, NodeAccess,
    PrimitiveTopology, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Scheme,
    ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const NUM_LINES: u32 = 20;

#[goldy::gpu]
struct Line {
    p1: [f32; 2],
    v1: [f32; 2],
    p2: [f32; 2],
    v2: [f32; 2],
    color_index: u32,
}

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut state = RenderState::new(None)?;
        while !state.capture_done() {
            state.render()?;
        }
        return Ok(());
    }

    println!("Goldy Bouncing Lines Example");
    println!("  Escape - Exit");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App::default();
    event_loop.run_app(&mut app)?;

    Ok(())
}

#[derive(Default)]
struct App {
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    device: Arc<goldy::Runtime>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    present: Option<Transaction>,
    scheme: Scheme,
    scene_rt: Lease<LeaseRenderTarget>,
    compute_pipeline: ComputePipeline,
    line_buffer: Buffer,
    render_shader: ShaderModule,
    render_pipeline: RenderPipeline,
    frame_count: u32,
    start_time: std::time::Instant,
}

impl RenderState {
    fn create_render_pipeline(
        device: &goldy::Runtime,
        render_shader: &ShaderModule,
        format: TextureFormat,
    ) -> Result<RenderPipeline> {
        common::render_pipeline(
            device,
            render_shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                topology: PrimitiveTopology::LineList,
                ..Default::default()
            },
        )
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn record_scheme(
        scheme: &mut Scheme,
        compute_pipeline: &ComputePipeline,
        render_pipeline: &RenderPipeline,
        line_buffer: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        scheme
            .node("update_lines", compute_pipeline)
            .with_parcel(line_buffer, NodeAccess::ReadWrite)
            .dispatch(NUM_LINES.div_ceil(64).max(1), 1, 1);

        let bg_color = Color {
            r: 0.05,
            g: 0.05,
            b: 0.1,
            a: 1.0,
        };

        let mut pass = scheme.render_pass("bouncing_lines", scene_rt, TargetLoad::Clear(bg_color));
        pass.with_parcel(line_buffer, NodeAccess::Read);
        pass.set_pipeline(render_pipeline);
        pass.draw(0..2, 0..NUM_LINES);
        pass.finish();
    }

    fn target(&self) -> (TextureFormat, u32, u32) {
        if let Some(surface) = &self.surface {
            let (width, height) = surface.size();
            (surface.format(), width, height)
        } else {
            let capture = self.capture.as_ref().expect("capture dump");
            let (width, height) = capture.size();
            (CaptureDump::format(), width, height)
        }
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn rerecord_scheme(&mut self) {
        let mut scheme = Scheme::new(&self.ctx);
        let (format, width, height) = self.target();
        if let Ok(rt) = self.ctx.lease_render_target(width.max(1), height.max(1), format, None) {
            self.scene_rt = rt;
            Self::record_scheme(
                &mut scheme,
                &self.compute_pipeline,
                &self.render_pipeline,
                &self.line_buffer,
                &self.scene_rt,
            );
            if let Ok(present) = Self::bind_frame(
                &mut scheme,
                &self.scene_rt,
                self.surface.as_ref(),
                self.readback.as_ref(),
            ) {
                self.present = present;
                self.scheme = scheme;
            }
        }
    }

    fn new(window: Option<Arc<Window>>) -> Result<Self> {
        let instance = Instance::new()?;
        let device = Arc::new(
            instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let compute_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/bouncing_lines_update.slang"),
            &[Line::GPU_TYPE],
        )?;
        let render_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/bouncing_lines_render.slang"),
            &[Line::GPU_TYPE],
        )?;

        let mut lines = Vec::with_capacity(NUM_LINES as usize);
        for idx in 0..NUM_LINES {
            let angle = (idx as f32 / NUM_LINES as f32) * std::f32::consts::PI * 2.0;
            lines.push(Line {
                p1: [angle.cos() * 0.3, angle.sin() * 0.3],
                v1: [0.01 * (idx as f32 * 0.7).cos(), 0.012 * (idx as f32 * 0.9).sin()],
                p2: [-angle.cos() * 0.3, -angle.sin() * 0.3],
                v2: [-0.011 * (idx as f32 * 1.1).cos(), 0.009 * (idx as f32 * 1.3).sin()],
                color_index: idx,
            });
        }

        let line_buffer = device.acquire_buffer_with_data(&lines, BufferKind::Scattered)?;

        let compute_pipeline = ComputePipeline::new(&device, &compute_shader)?;
        let render_pipeline = Self::create_render_pipeline(&device, &render_shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_scheme(
            &mut scheme,
            &compute_pipeline,
            &render_pipeline,
            &line_buffer,
            &scene_rt,
        );
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        println!("Created bouncing lines with {} lines", NUM_LINES);

        Ok(Self {
            window,
            device,
            ctx,
            surface,
            capture,
            readback,
            present,
            scheme,
            scene_rt,
            compute_pipeline,
            line_buffer,
            render_shader,
            render_pipeline,
            frame_count: 0,
            start_time: std::time::Instant::now(),
        })
    }

    fn render(&mut self) -> Result<()> {
        self.frame_count += 1;

        let mut submission = self.scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        if let Some(window) = &self.window {
            window.request_redraw();
        }
        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Bouncing Lines", 800, 600))
                    .expect("Failed to create window"),
            );

            match RenderState::new(Some(window.clone())) {
                Ok(mut state) => {
                    if let Err(e) = state.render() {
                        tracing::error!("First frame error: {e}");
                    }
                    common::reveal_window(&window);
                    self.state = Some(state);
                    window.request_redraw();
                }
                Err(e) => {
                    tracing::error!("Failed to create render state: {}", e);
                    event_loop.exit();
                }
            }
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => {
                event_loop.exit();
            }
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::Resized(size) => {
                if let Some(state) = &mut self.state {
                    if size.width > 0 && size.height > 0 {
                        let Some(surface) = state.surface.as_ref() else {
                            return;
                        };
                        let (prev_w, prev_h) = surface.size();
                        if size.width == prev_w && size.height == prev_h {
                            return;
                        }
                        let _ = surface.resize(size.width, size.height);
                        let format = surface.format();
                        if let Ok(pipeline) =
                            RenderState::create_render_pipeline(&state.device, &state.render_shader, format)
                        {
                            state.render_pipeline = pipeline;
                        }
                        state.rerecord_scheme();
                    }
                }
            }
            WindowEvent::RedrawRequested => {
                if let Some(state) = &mut self.state {
                    if let Err(e) = state.render() {
                        tracing::error!("Render error: {}", e);
                    }
                }
            }
            _ => {}
        }
    }
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/bouncing_lines_update.slang:

// Bouncing lines compute shader
// Updates line endpoint positions with wall bouncing physics

import goldy_exp;


static const uint NUM_LINES = 20;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Line> LINES, ThreadId id) {
    uint idx = id.x;
    
    if (idx >= NUM_LINES) return;
    
    Line line = LINES[idx];
    
    // Update positions
    line.p1 += line.v1;
    line.p2 += line.v2;
    
    // Bounce endpoint 1 off walls
    if (line.p1.x < -1.0 || line.p1.x > 1.0) {
        line.v1.x = -line.v1.x;
        line.p1.x = clamp(line.p1.x, -1.0, 1.0);
    }
    if (line.p1.y < -1.0 || line.p1.y > 1.0) {
        line.v1.y = -line.v1.y;
        line.p1.y = clamp(line.p1.y, -1.0, 1.0);
    }
    
    // Bounce endpoint 2 off walls
    if (line.p2.x < -1.0 || line.p2.x > 1.0) {
        line.v2.x = -line.v2.x;
        line.p2.x = clamp(line.p2.x, -1.0, 1.0);
    }
    if (line.p2.y < -1.0 || line.p2.y > 1.0) {
        line.v2.y = -line.v2.y;
        line.p2.y = clamp(line.p2.y, -1.0, 1.0);
    }
    
    LINES[idx] = line;
}

shaders/bouncing_lines_render.slang:

// Bouncing lines rendering shader
// Visualizes lines using instancing with LineList topology

import goldy_exp;


struct VSOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

static const float4 colors[6] = {
    float4(1.0, 0.0, 0.0, 1.0),  // Red
    float4(0.0, 1.0, 0.0, 1.0),  // Green
    float4(0.0, 0.0, 1.0, 1.0),  // Blue
    float4(1.0, 1.0, 0.0, 1.0),  // Yellow
    float4(1.0, 0.0, 1.0, 1.0),  // Magenta
    float4(0.0, 1.0, 1.0, 1.0),  // Cyan
};

[goldy_vertex]
VSOutput vs_main(Scattered<Line> LINES, VertexId vertexID, InstanceId instanceID) {
    VSOutput output;
    
    Line line = LINES[instanceID.value];
    
    float2 pos = (vertexID.value == 0) ? line.p1 : line.p2;
    
    output.position = float4(pos, 0.0, 1.0);
    output.color = colors[line.color_index % 6];
    
    return output;
}

[goldy_fragment]
float4 fs_main(VSOutput input) : SV_Target {
    return input.color;
}

waveform

An animated waveform visualizer built from a LINE_STRIP whose vertices are recomputed on the CPU and uploaded each frame.

cargo run --features examples --example waveform

What it demonstrates

  • LINE_STRIP topology
  • Per-frame vertex buffer deposits

Source

examples/waveform.rs:

//! Waveform example - animated audio waveform visualizer.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example waveform`

use goldy::{
    Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease, LeaseRenderTarget,
    MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, Vertex2D,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const NUM_SAMPLES: usize = 200;
const NUM_CHANNELS: usize = 4;

fn generate_waveform(time: f32, channel: usize) -> Vec<f32> {
    let freq = 1.0 + channel as f32 * 0.5;
    let phase = channel as f32 * 0.7;

    (0..NUM_SAMPLES)
        .map(|i| {
            let x = i as f32 / NUM_SAMPLES as f32 * 6.0;
            let mut y = 0.0;

            // Superposition of different frequencies
            y += (x * freq + time * 2.0 + phase).sin() * 0.3;
            y += (x * freq * 2.3 + time * 1.7 + phase).sin() * 0.2;
            y += (x * freq * 3.7 + time * 0.9 + phase).cos() * 0.15;
            y += (x * freq * 5.1 + time * 2.3 + phase).sin() * 0.1;

            // Add some noise
            let noise = ((i as f32 * 1234.5 + time * 100.0).sin() * 43_758.547_f32).fract() - 0.5;
            y += noise * 0.05;

            y.clamp(-1.0, 1.0)
        })
        .collect()
}

fn waveform_to_vertices(samples: &[f32], y_offset: f32, color: Color) -> Vec<Vertex2D> {
    let scale_y = 0.15;
    samples
        .iter()
        .enumerate()
        .map(|(i, &sample)| {
            let x = (i as f32 / (NUM_SAMPLES - 1) as f32) * 1.9 - 0.95;
            let y = y_offset + sample * scale_y;
            Vertex2D::new(x, y, color)
        })
        .collect()
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    channel_parcels: Option<[Buffer; NUM_CHANNELS]>,
    upload_scheme: Option<Scheme>,
    channel_deposits: Option<[DepositTransaction; NUM_CHANNELS]>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            start_time: Instant::now(),
            frame_count: 0,
            channel_parcels: None,
            upload_scheme: None,
            channel_deposits: None,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: Vertex2D::layout(),
                topology: PrimitiveTopology::LineStrip,
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        channel_parcels: &[Buffer; NUM_CHANNELS],
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass(
            "waveform",
            scene_rt,
            TargetLoad::Clear(Color {
                r: 0.02,
                g: 0.02,
                b: 0.08,
                a: 1.0,
            }),
        );
        for parcel in channel_parcels {
            pass.with_parcel(parcel, NodeAccess::Read);
        }

        pass.set_pipeline(pipeline);
        for parcel in channel_parcels {
            pass.set_vertex_buffer(0, parcel);
            pass.draw(0..NUM_SAMPLES as u32, 0..1);
        }
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, goldy::shader::builtins::VERTEX_COLOR_2D)?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let channel_parcels = std::array::from_fn(|_| {
            device
                .acquire_buffer_sized::<Vertex2D>(NUM_SAMPLES as u64, BufferKind::Scattered, BufferFlags::empty())
                .expect("waveform channel parcel")
        });

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &channel_parcels, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        let ctx = self.ctx.as_ref().unwrap();
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.channel_parcels = Some(channel_parcels);
        let channel_parcels = self.channel_parcels.as_ref().unwrap();
        let mut upload_scheme = Scheme::new(ctx);
        let memory = MemoryExchange::new(ctx);
        let channel_capacity = channel_parcels[0].byte_size();
        let channel_deposits = std::array::from_fn(|ch| {
            memory
                .bind_deposit(
                    &mut upload_scheme,
                    DepositTarget::buffer(&channel_parcels[ch], channel_capacity),
                )
                .expect("bind channel deposit")
        });
        self.upload_scheme = Some(upload_scheme);
        self.channel_deposits = Some(channel_deposits);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();
        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());

        let colors = [
            Color {
                r: 1.0,
                g: 0.3,
                b: 0.3,
                a: 1.0,
            },
            Color {
                r: 0.3,
                g: 1.0,
                b: 0.3,
                a: 1.0,
            },
            Color {
                r: 0.3,
                g: 0.5,
                b: 1.0,
                a: 1.0,
            },
            Color {
                r: 1.0,
                g: 0.8,
                b: 0.2,
                a: 1.0,
            },
        ];
        let y_offsets = [0.6, 0.2, -0.2, -0.6];

        let upload = self.upload_scheme.as_mut().unwrap();
        let channel_deposits = self.channel_deposits.as_ref().unwrap();
        for ch in 0..NUM_CHANNELS {
            let samples = generate_waveform(time, ch);
            let vertices = waveform_to_vertices(&samples, y_offsets[ch], colors[ch]);
            (&channel_deposits[ch] << vertices.as_slice())?;
        }
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(channel_parcels)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.channel_parcels.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, channel_parcels, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Waveform Visualizer (Scheme + Present)",
                        1024,
                        600,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Waveform Example - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/vertex_color_2d.slang:

// Simple 2D vertex + fragment shader for colored vertices.
// Used by: triangle, particles, starfield, bouncing_lines, spinning_cube, instancing, waveform

struct VertexInput {
    float2 position : POSITION;
    float4 color : COLOR;
};

struct VertexOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[goldy_vertex]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.color = input.color;
    return output;
}

[goldy_fragment]
float4 fs_main(VertexOutput input) : SV_Target {
    return input.color;
}

plasma

The classic demoscene plasma effect: layered sine fields evaluated per pixel in the fragment shader, with time as the only input.

cargo run --features examples --example plasma

What it demonstrates

  • Vertex-less fullscreen fragment effect
  • Single time uniform driving the whole image

Source

examples/plasma.rs:

//! Plasma example - classic demoscene plasma effect.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example plasma`

use goldy::{
    shaders, Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct TimeUniforms {
    time: f32,
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    uniform: Option<Buffer>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    uniform_deposit: Option<DepositTransaction>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            uniform: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            uniform_deposit: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        uniform: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("plasma", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(uniform, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.draw_fullscreen();
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang_with_gpu_types(&device, shaders::PLASMA, &[TimeUniforms::GPU_TYPE])?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let uniform = device.acquire_buffer_sized::<TimeUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &uniform, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<TimeUniforms>(&uniform, 1),
        )?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.uniform = Some(uniform);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.upload_scheme = Some(upload_scheme);
        self.uniform_deposit = Some(uniform_deposit);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());
        let uniforms = TimeUniforms { time };
        let upload = self.upload_scheme.as_mut().unwrap();
        (self.uniform_deposit.as_ref().unwrap() << &uniforms)?;
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(uniform)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.uniform.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, uniform, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Plasma Effect (Scheme + Present)",
                        800,
                        600,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Plasma Example - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/plasma.slang:

// Classic demoscene plasma effect
// Uses vertex-less fullscreen triangle and time uniform

import goldy_exp;


[goldy_vertex]
FullscreenVarying vs_main(VertexId vertex_id) {
    return vs_fullscreen_triangle(vertex_id.value);
}

[goldy_fragment]
float4 fs_main(TimeUniforms uniforms, FullscreenVarying input) : SV_Target {
    float2 uv = scale_uv(input.uv, 4.0);
    float t = uniforms.time;
    
    // Classic plasma formula
    float v = sin(uv.x + t);
    v += sin(uv.y + t);
    v += sin(uv.x + uv.y + t);
    
    float cx = uv.x + 0.5 * sin(t / 3.0);
    float cy = uv.y + 0.5 * cos(t / 2.0);
    v += sin(sqrt(cx * cx + cy * cy + 1.0) + t);
    
    v = v / 2.0;
    
    return float4(rainbow(v), 1.0);
}

tunnel

A flight through an endless tunnel, produced by converting screen coordinates to polar coordinates and scrolling a procedural texture along them.

cargo run --features examples --example tunnel

What it demonstrates

  • Polar-coordinate screen-space effects
  • Vertex-less fullscreen rendering

Source

examples/tunnel.rs:

//! Tunnel example - classic demoscene tunnel effect.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example tunnel`

use goldy::{
    shaders, Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct TimeUniforms {
    time: f32,
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    uniform: Option<Buffer>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    uniform_deposit: Option<DepositTransaction>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            uniform: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            uniform_deposit: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        uniform: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("tunnel", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(uniform, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.draw_fullscreen();
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang_with_gpu_types(&device, shaders::TUNNEL, &[TimeUniforms::GPU_TYPE])?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let uniform = device.acquire_buffer_sized::<TimeUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &uniform, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<TimeUniforms>(&uniform, 1),
        )?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.uniform = Some(uniform);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.upload_scheme = Some(upload_scheme);
        self.uniform_deposit = Some(uniform_deposit);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());
        let uniforms = TimeUniforms { time };
        let upload = self.upload_scheme.as_mut().unwrap();
        (self.uniform_deposit.as_ref().unwrap() << &uniforms)?;
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(uniform)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.uniform.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, uniform, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Tunnel Effect (Scheme + Present)",
                        800,
                        800,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Tunnel Example - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/tunnel.slang:

// Classic demoscene tunnel effect
// Uses goldy_exp module for vertex format utilities

import goldy_exp;


[goldy_vertex]
FullscreenVarying vs_main(VertexId vertex_id) {
    return vs_fullscreen_triangle(vertex_id.value);
}

[goldy_fragment]
float4 fs_main(TimeUniforms uniforms, FullscreenVarying input) : SV_Target {
    float2 uv = (input.uv - 0.5) * 2.0;
    float t = uniforms.time;
    
    // Polar coordinates
    float dist = length(uv);
    float angle = atan2(uv.y, uv.x);
    
    // Tunnel coordinates
    float tunnel_depth = 1.0 / (dist + 0.1);
    float tunnel_angle = angle / 3.14159 + t * 0.2;
    
    // Animated texture coordinates
    float tx = tunnel_angle * 4.0;
    float ty = tunnel_depth - t * 2.0;
    
    // Checkerboard pattern
    float checker = floor(tx) + floor(ty);
    bool is_white = fmod(checker, 2.0) == 0.0;
    
    // Color based on depth and checker
    float depth_color = 1.0 - dist * 0.5;
    float3 color;
    
    if (is_white) {
        color = float3(0.8, 0.2, 0.4) * depth_color;
    } else {
        color = float3(0.2, 0.4, 0.8) * depth_color;
    }
    
    // Add glow at center
    color += float3(0.3, 0.5, 1.0) * (1.0 - dist) * (1.0 - dist);
    
    // Fog at edges
    color *= 1.0 - dist * 0.3;
    
    return float4(color, 1.0);
}

metaballs

Organic blobs rendered by summing an inverse-square field from several moving centres and thresholding it in the fragment shader.

cargo run --features examples --example metaballs

What it demonstrates

  • Scalar-field rendering in a fragment shader
  • Animated uniform arrays

Source

examples/metaballs.rs:

//! Metaballs example - organic blob simulation.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example metaballs`

use goldy::{
    shaders, Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct TimeUniforms {
    time: f32,
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    uniform: Option<Buffer>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    uniform_deposit: Option<DepositTransaction>,
    start_time: Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            uniform: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            uniform_deposit: None,
            start_time: Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        uniform: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("metaballs", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(uniform, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.draw_fullscreen();
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang_with_gpu_types(&device, shaders::METABALLS, &[TimeUniforms::GPU_TYPE])?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let uniform = device.acquire_buffer_sized::<TimeUniforms>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &uniform, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<TimeUniforms>(&uniform, 1),
        )?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.uniform = Some(uniform);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.upload_scheme = Some(upload_scheme);
        self.uniform_deposit = Some(uniform_deposit);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();

        let time = self
            .capture
            .as_ref()
            .map(CaptureDump::time)
            .unwrap_or_else(|| self.start_time.elapsed().as_secs_f32());
        let uniforms = TimeUniforms { time };
        let upload = self.upload_scheme.as_mut().unwrap();
        (self.uniform_deposit.as_ref().unwrap() << &uniforms)?;
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(uniform)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.uniform.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, uniform, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Metaballs (Scheme + Present)", 800, 800))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if matches!(event.logical_key, Key::Named(NamedKey::Escape)) {
                    event_loop.exit();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                self.window.as_ref().unwrap().request_redraw();
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Metaballs Example - Press Escape to exit");
    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/metaballs.slang:

// Animated metaballs effect
// Uses goldy_exp module for vertex format utilities

import goldy_exp;


[goldy_vertex]
FullscreenVarying vs_main(VertexId vertex_id) {
    return vs_fullscreen_triangle(vertex_id.value);
}

float metaball(float2 p, float2 center, float radius) {
    float d = distance(p, center);
    return radius / (d * d + 0.001);
}

[goldy_fragment]
float4 fs_main(TimeUniforms uniforms, FullscreenVarying input) : SV_Target {
    float2 uv = (input.uv - 0.5) * 2.0;
    float t = uniforms.time;
    
    // Moving metaball centers
    float2 c1 = float2(sin(t * 1.1) * 0.5, cos(t * 0.9) * 0.5);
    float2 c2 = float2(sin(t * 0.8 + 1.0) * 0.6, cos(t * 1.2 + 2.0) * 0.4);
    float2 c3 = float2(sin(t * 1.3 + 2.0) * 0.4, cos(t * 0.7 + 1.0) * 0.6);
    float2 c4 = float2(cos(t * 0.9) * 0.5, sin(t * 1.1 + 3.0) * 0.5);
    float2 c5 = float2(cos(t * 1.0 + 1.5) * 0.3, sin(t * 0.8 + 0.5) * 0.7);
    
    // Sum metaball influences
    float v = 0.0;
    v += metaball(uv, c1, 0.15);
    v += metaball(uv, c2, 0.12);
    v += metaball(uv, c3, 0.18);
    v += metaball(uv, c4, 0.10);
    v += metaball(uv, c5, 0.14);
    
    // Threshold and color
    float threshold = 1.0;
    if (v > threshold) {
        float intensity = (v - threshold) / 2.0;
        float r = 0.2 + intensity * 0.3;
        float g = 0.5 + intensity * 0.4;
        float b = 0.8 + intensity * 0.2;
        return float4(r, g, b, 1.0);
    } else {
        float glow = v * 0.3;
        return float4(glow * 0.2, glow * 0.3, glow * 0.5, 1.0);
    }
}

mandelbrot

An interactive Mandelbrot explorer. Pan and zoom update a uniform that the fragment shader uses as its complex-plane window, so navigation costs nothing but a deposit.

cargo run --features examples --example mandelbrot

What it demonstrates

  • Interactive uniform updates from keyboard input
  • Iteration-count colouring in a fragment shader

Controls

KeyAction
Arrow keysPan
+ / =Zoom in
-Zoom out
RReset view
EscapeExit

Source

examples/mandelbrot.rs:

//! Mandelbrot example - interactive fractal explorer.
//!
//! Demonstrates retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example mandelbrot`

use goldy::{
    shaders, Buffer, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease, LeaseRenderTarget,
    MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Scheme,
    ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat, Transaction,
};
use std::ops::Shr;
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

#[goldy::gpu]
struct ViewUniforms {
    center: [f32; 2],
    zoom: f32,
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    uniform: Option<Buffer>,
    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,
    upload_scheme: Option<Scheme>,
    uniform_deposit: Option<DepositTransaction>,
    center: [f32; 2],
    zoom: f32,
    start_time: std::time::Instant,
    frame_count: u32,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            uniform: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            upload_scheme: None,
            uniform_deposit: None,
            center: [-0.5, 0.0],
            zoom: 1.0,
            start_time: std::time::Instant::now(),
            frame_count: 0,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: goldy::VertexBufferLayout::empty(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        uniform: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("mandelbrot", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(uniform, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.draw_fullscreen();
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang_with_gpu_types(&device, shaders::MANDELBROT, &[ViewUniforms::GPU_TYPE])?;

        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let uniform = device.acquire_buffer_with_data(
            &[ViewUniforms {
                center: self.center,
                zoom: self.zoom,
            }],
            BufferKind::Broadcast,
        )?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &uniform, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let uniform_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<ViewUniforms>(&uniform, 1),
        )?;

        self.ctx = Some(ctx);
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.uniform = Some(uniform);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.upload_scheme = Some(upload_scheme);
        self.uniform_deposit = Some(uniform_deposit);
        Ok(())
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let scheme = self.scheme.as_mut().unwrap();

        let uniforms = ViewUniforms {
            center: self.center,
            zoom: self.zoom,
        };
        let upload = self.upload_scheme.as_mut().unwrap();
        (self.uniform_deposit.as_ref().unwrap() << &uniforms)?;
        upload.submit()?;

        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(uniform)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.uniform.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, uniform, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window(
                        "Goldy - Mandelbrot (Scheme + Present, Arrows=pan, +/-=zoom, R=reset)",
                        800,
                        800,
                    ))
                    .unwrap(),
            );
            self.window = Some(window.clone());
            self.init_gpu(Some(window.as_ref())).unwrap();
            if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                let pan = 0.1 / self.zoom;
                match event.logical_key {
                    Key::Named(NamedKey::Escape) => event_loop.exit(),
                    Key::Named(NamedKey::ArrowUp) => self.center[1] += pan,
                    Key::Named(NamedKey::ArrowDown) => self.center[1] -= pan,
                    Key::Named(NamedKey::ArrowLeft) => self.center[0] -= pan,
                    Key::Named(NamedKey::ArrowRight) => self.center[0] += pan,
                    Key::Character(ref c) if c == "=" || c == "+" => self.zoom *= 1.5,
                    Key::Character(ref c) if c == "-" => self.zoom /= 1.5,
                    Key::Character(ref c) if c == "r" || c == "R" => {
                        self.center = [-0.5, 0.0];
                        self.zoom = 1.0;
                    }
                    _ => {}
                }
                if let Some(w) = &self.window {
                    w.request_redraw();
                }
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Mandelbrot Example");
    println!("  Arrows - Pan");
    println!("  +/- - Zoom in/out");
    println!("  R - Reset view");
    println!("  Escape - Exit");
    let event_loop = EventLoop::new()?;
    // `Wait` idles until input arrives, which would also idle straight past a
    // configured run limit, so poll when one is set.
    event_loop.set_control_flow(if common::run_limit_secs().is_some() {
        ControlFlow::Poll
    } else {
        ControlFlow::Wait
    });
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/mandelbrot.slang:

// Interactive Mandelbrot fractal explorer
// Uses goldy_exp module for vertex format and color utilities

import goldy_exp;


[goldy_vertex]
FullscreenVarying vs_main(VertexId vertex_id) {
    FullscreenVarying output;
    
    float2 uv = float2(
        float((vertex_id.value << 1u) & 2u),
        float(vertex_id.value & 2u)
    );
    
    output.position = float4(uv * 2.0 - 1.0, 0.0, 1.0);
    output.uv = uv;
    
    return output;
}

[goldy_fragment]
float4 fs_main(ViewUniforms view, FullscreenVarying input) : SV_Target {
    float2 c = view.center + (input.uv - 0.5) * 4.0 / view.zoom;
    float2 z = float2(0.0, 0.0);
    uint i = 0;
    const uint max_iter = 256;
    
    for (i = 0; i < max_iter; i++) {
        if (dot(z, z) > 4.0) break;
        z = float2(z.x * z.x - z.y * z.y, 2.0 * z.x * z.y) + c;
    }
    
    if (i >= max_iter) {
        return float4(0.0, 0.0, 0.0, 1.0);
    }
    
    float t = float(i) / float(max_iter);
    float3 color = palette(t, float3(5.0, 7.0, 11.0), float3(0.0, 1.0, 2.0));
    
    return float4(color, 1.0);
}

digital_clock

A seven-segment clock. Segment geometry is generated on the CPU in examples/digital_clock_shared.rs and drawn as coloured triangles each frame.

cargo run --features examples --example digital_clock

What it demonstrates

  • Dynamic CPU-generated geometry
  • Sharing helper modules between examples

Controls

KeyAction
SpacePause / resume
CCycle colour
EscapeExit

Source

examples/digital_clock.rs:

//! Digital Clock example - render an animated 7-segment clock in a window.
//!
//! Uses retained scheme with offscreen render pass → copy-to-present.
//!
//! Run with: `cargo run --example digital_clock`

mod digital_clock_shared;

use digital_clock_shared::{generate_clock_vertices, ClockState, ClockVertex, TimeData};
use goldy::{
    Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease, LeaseRenderTarget,
    MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions, RuntimeDescriptor, Scheme,
    ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

/// Upper bound on seven-segment clock vertices (8 glyphs × 7 segments × 6 verts).
const MAX_CLOCK_VERTICES: usize = 384;

const SHADER_SOURCE: &str = r#"
struct VertexInput {
    float2 position : POSITION;
    float4 color : COLOR;
};

struct VertexOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

[shader("vertex")]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.color = input.color;
    return output;
}

[shader("fragment")]
float4 fs_main(VertexOutput input) : SV_Target {
    return input.color;
}
"#;

fn clock_vertex_layout() -> VertexBufferLayout {
    ClockVertex::GPU_TYPE
        .vertex_buffer_layout()
        .expect("clock vertex layout")
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    pipeline: Option<RenderPipeline>,
    shader: Option<ShaderModule>,
    vertex_parcel: Option<Buffer>,
    upload_scheme: Option<Scheme>,
    vertex_deposit: Option<DepositTransaction>,

    window: Option<Arc<Window>>,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scene_rt: Option<Lease<LeaseRenderTarget>>,
    scheme: Option<Scheme>,

    start_time: Instant,
    perf_start: Instant,
    frame_count: u32,
    clock_state: ClockState,
    recorded_vertex_count: u32,
    recorded_bg_color: Color,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        let instance = Instance::new()?;
        Ok(Self {
            instance,
            ctx: None,
            device: None,
            pipeline: None,
            shader: None,
            window: None,
            surface: None,
            present: None,
            capture: None,
            readback: None,
            scene_rt: None,
            scheme: None,
            start_time: Instant::now(),
            perf_start: Instant::now(),
            frame_count: 0,
            clock_state: ClockState::default(),
            vertex_parcel: None,
            upload_scheme: None,
            vertex_deposit: None,
            recorded_vertex_count: 0,
            recorded_bg_color: Color::BLACK,
        })
    }

    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: clock_vertex_layout(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        vertex_parcel: &Buffer,
        vertex_count: u32,
        bg_color: Color,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        let mut pass = scheme.render_pass("digital_clock", scene_rt, TargetLoad::Clear(bg_color));
        pass.with_parcel(vertex_parcel, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, vertex_parcel);
        pass.draw(0..vertex_count, 0..1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn output_size(&self) -> Option<(TextureFormat, u32, u32)> {
        if let Some(surface) = self.surface.as_ref() {
            let (width, height) = surface.size();
            Some((surface.format(), width, height))
        } else {
            let capture = self.capture.as_ref()?;
            Some((CaptureDump::format(), capture.width(), capture.height()))
        }
    }

    fn rerecord_scheme_if_needed(&mut self, vertex_count: u32, bg_color: Color) {
        if vertex_count == self.recorded_vertex_count && bg_color == self.recorded_bg_color {
            return;
        }
        let Some((format, width, height)) = self.output_size() else {
            return;
        };
        if let (Some(ctx), Some(pipeline), Some(vertex_parcel)) =
            (self.ctx.as_ref(), self.pipeline.as_ref(), self.vertex_parcel.as_ref())
        {
            let mut scheme = Scheme::new(ctx);
            if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                Self::record_pass(&mut scheme, pipeline, vertex_parcel, vertex_count, bg_color, &rt);
                if let Ok(present) = Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref()) {
                    self.present = present;
                    self.scheme = Some(scheme);
                    self.recorded_vertex_count = vertex_count;
                    self.recorded_bg_color = bg_color;
                    self.scene_rt = Some(rt);
                }
            }
        }
    }

    fn init_gpu(&mut self, window: Option<&Window>) -> anyhow::Result<()> {
        let device = Arc::new(
            self.instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let shader = ShaderModule::from_slang(&device, SHADER_SOURCE)?;
        let pipeline = Self::create_pipeline(&device, &shader, format)?;

        let vertex_parcel = device.acquire_buffer_sized::<ClockVertex>(
            MAX_CLOCK_VERTICES as u64,
            BufferKind::Scattered,
            BufferFlags::empty(),
        )?;

        let bg_color = self.clock_state.background_color();
        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &vertex_parcel, 1, bg_color, &scene_rt);
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        self.ctx = Some(ctx);
        let ctx = self.ctx.as_ref().unwrap();
        self.device = Some(device);
        self.shader = Some(shader);
        self.pipeline = Some(pipeline);
        self.vertex_parcel = Some(vertex_parcel);
        let vertex_parcel = self.vertex_parcel.as_ref().unwrap();
        let mut upload_scheme = Scheme::new(ctx);
        let vertex_deposit = MemoryExchange::new(ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<ClockVertex>(vertex_parcel, MAX_CLOCK_VERTICES as u64),
        )?;
        self.upload_scheme = Some(upload_scheme);
        self.vertex_deposit = Some(vertex_deposit);
        self.surface = surface;
        self.present = present;
        self.capture = capture;
        self.readback = readback;
        self.scene_rt = Some(scene_rt);
        self.scheme = Some(scheme);
        self.recorded_vertex_count = 1;
        self.recorded_bg_color = bg_color;
        Ok(())
    }

    fn elapsed_secs(&self) -> u64 {
        if let Some(capture) = self.capture.as_ref() {
            return capture.time() as u64;
        }
        if self.clock_state.paused {
            self.clock_state.accumulated_secs
        } else {
            self.start_time.elapsed().as_secs() + self.clock_state.accumulated_secs
        }
    }

    fn toggle_pause(&mut self) {
        let current = self.elapsed_secs();
        if self.clock_state.paused {
            self.start_time = Instant::now();
        }
        self.clock_state.toggle_pause(current);
    }

    fn render_frame(&mut self) -> anyhow::Result<()> {
        self.frame_count += 1;

        let (width, height) = if let Some(window) = self.window.as_ref() {
            let size = window.inner_size();
            (size.width, size.height)
        } else {
            self.capture.as_ref().unwrap().size()
        };

        if width == 0 || height == 0 {
            return Ok(());
        }

        let elapsed = self.elapsed_secs();
        let time = TimeData::from_elapsed_secs(elapsed);
        let color = self.clock_state.color();
        let bg_color = self.clock_state.background_color();
        let vertices = generate_clock_vertices(time, color, width, height);
        let vertex_count = vertices.len() as u32;

        self.rerecord_scheme_if_needed(vertex_count, bg_color);

        let upload = self.upload_scheme.as_mut().unwrap();
        (self.vertex_deposit.as_ref().unwrap() << vertices.as_slice())?;
        upload.submit()?;

        let scheme = self.scheme.as_mut().unwrap();
        let mut submission = scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn handle_resize(&mut self, new_size: winit::dpi::PhysicalSize<u32>) {
        if new_size.width == 0 || new_size.height == 0 {
            return;
        }
        let Some(surface) = self.surface.as_mut() else {
            return;
        };
        let _ = surface.resize(new_size.width, new_size.height);
        let format = surface.format();
        let (width, height) = surface.size();
        if let (Some(ctx), Some(device), Some(shader), Some(vertex_parcel)) = (
            self.ctx.as_ref(),
            self.device.as_ref(),
            self.shader.as_ref(),
            self.vertex_parcel.as_ref(),
        ) {
            if let Ok(pipeline) = Self::create_pipeline(device, shader, format) {
                self.pipeline = Some(pipeline);
                if let Some(pipeline) = self.pipeline.as_ref() {
                    let bg_color = self.clock_state.background_color();
                    let vertex_count = self.recorded_vertex_count.max(1);
                    let mut scheme = Scheme::new(ctx);
                    if let Ok(rt) = ctx.lease_render_target(width.max(1), height.max(1), format, None) {
                        Self::record_pass(&mut scheme, pipeline, vertex_parcel, vertex_count, bg_color, &rt);
                        if let Ok(present) =
                            Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref())
                        {
                            self.present = present;
                            self.scheme = Some(scheme);
                            self.recorded_vertex_count = vertex_count;
                            self.recorded_bg_color = bg_color;
                            self.scene_rt = Some(rt);
                        }
                    }
                }
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.perf_start.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.window.is_none() {
            let attrs = common::hidden_window(
                "Goldy - Clock (Scheme + Present, Space: pause, Click: color)",
                1280,
                720,
            );

            let window = Arc::new(event_loop.create_window(attrs).unwrap());
            self.window = Some(window.clone());

            if let Err(e) = self.init_gpu(Some(window.as_ref())) {
                tracing::error!("Failed to initialize GPU: {}", e);
            } else if let Err(e) = self.render_frame() {
                tracing::error!("First frame error: {e}");
            }
            common::reveal_window(&window);
            window.request_redraw();
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.perf_start);
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => {
                event_loop.exit();
            }
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => match event.logical_key {
                Key::Named(NamedKey::Escape) => event_loop.exit(),
                Key::Named(NamedKey::Space) => self.toggle_pause(),
                Key::Character(ref c) if c == "c" || c == "C" => self.clock_state.next_color(),
                _ => {}
            },
            WindowEvent::MouseInput { state, .. } if state.is_pressed() => {
                self.clock_state.next_color();
            }
            WindowEvent::RedrawRequested => {
                if let Err(e) = self.render_frame() {
                    tracing::error!("Render error: {}", e);
                }
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            WindowEvent::Resized(new_size) => {
                self.handle_resize(new_size);
                if let Some(window) = &self.window {
                    window.request_redraw();
                }
            }
            _ => {}
        }
    }
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut app = App::new()?;
        app.init_gpu(None)?;
        while !app.capture_done() {
            app.render_frame()?;
        }
        return Ok(());
    }

    println!("Goldy Clock Example (retained scheme)");
    println!("==================================================================");
    println!("Controls:");
    println!("  Space - Toggle pause");
    println!("  Click - Change color");
    println!("  Escape - Exit\n");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App::new()?;
    event_loop.run_app(&mut app)?;

    Ok(())
}

The example pulls in examples/common.rs, examples/digital_clock_shared.rs — see Shared Helpers.

The Slang source is inline in the example above.

starfield

A 3D starfield flying towards the viewer. A compute pass advances and recycles stars; the raster pass scales each point by its depth.

cargo run --features examples --example starfield

What it demonstrates

  • Compute-driven particle recycling
  • Depth-scaled point rendering

Source

examples/starfield.rs:

//! Starfield example - classic 3D starfield flying through space.
//!
//! Demonstrates retained scheme with compute dispatch → offscreen render → copy-to-present.
//!
//! Run with: `cargo run --example starfield`

use anyhow::Result;
use goldy::{
    Buffer, BufferFlags, BufferKind, Color, ComputePipeline, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc,
    RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad,
    Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const NUM_STARS: u32 = 500;

const STAR_TYPE_NORMAL: f32 = 0.0;
const STAR_TYPE_GALAXY: f32 = 1.0;
const STAR_TYPE_QUASAR: f32 = 2.0;
const STAR_TYPE_WHITE_DWARF: f32 = 3.0;

#[goldy::gpu]
struct Star {
    x: f32,
    y: f32,
    z: f32,
    star_type: f32,
}

#[goldy::gpu]
struct StarfieldParams {
    speed: f32,
    frame: f32,
}

static mut SEED: u32 = 12345;
fn rand_f32() -> f32 {
    unsafe {
        SEED = SEED.wrapping_mul(1103515245).wrapping_add(12345);
        SEED as f32 / u32::MAX as f32
    }
}

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut state = RenderState::new(None)?;
        while !state.capture_done() {
            state.render()?;
        }
        return Ok(());
    }

    println!("Goldy Starfield Example");
    println!("  Up/Down - Change speed");
    println!("  Escape - Exit");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App::default();
    event_loop.run_app(&mut app)?;

    Ok(())
}

#[derive(Default)]
struct App {
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    device: Arc<goldy::Runtime>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    present: Option<Transaction>,
    scheme: Scheme,
    scene_rt: Lease<LeaseRenderTarget>,
    compute_pipeline: ComputePipeline,
    render_shader: ShaderModule,
    render_pipeline: RenderPipeline,
    star_buffer: Buffer,
    params_buffer: Buffer,
    upload_scheme: Scheme,
    params_deposit: DepositTransaction,
    speed: f32,
    frame_count: f32,
    start_time: std::time::Instant,
}

impl RenderState {
    fn create_render_pipeline(
        device: &goldy::Runtime,
        render_shader: &ShaderModule,
        format: TextureFormat,
    ) -> Result<RenderPipeline> {
        common::render_pipeline(
            device,
            render_shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                topology: PrimitiveTopology::TriangleList,
                ..Default::default()
            },
        )
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn record_scheme(
        scheme: &mut Scheme,
        compute_pipeline: &ComputePipeline,
        render_pipeline: &RenderPipeline,
        star_buffer: &Buffer,
        params_buffer: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
    ) {
        scheme
            .node("update_stars", compute_pipeline)
            .with_parcel(star_buffer, NodeAccess::ReadWrite)
            .with_parcel(params_buffer, NodeAccess::Read)
            .dispatch(NUM_STARS.div_ceil(64), 1, 1);

        let mut pass = scheme.render_pass("starfield", scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(star_buffer, NodeAccess::Read);
        pass.set_pipeline(render_pipeline);
        pass.draw(0..6, 0..NUM_STARS);
        pass.finish();
    }

    fn target(&self) -> (TextureFormat, u32, u32) {
        if let Some(surface) = &self.surface {
            let (width, height) = surface.size();
            (surface.format(), width, height)
        } else {
            let capture = self.capture.as_ref().expect("capture dump");
            let (width, height) = capture.size();
            (CaptureDump::format(), width, height)
        }
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn rerecord_scheme(&mut self) {
        let mut scheme = Scheme::new(&self.ctx);
        let (format, width, height) = self.target();
        if let Ok(rt) = self.ctx.lease_render_target(width.max(1), height.max(1), format, None) {
            self.scene_rt = rt;
            Self::record_scheme(
                &mut scheme,
                &self.compute_pipeline,
                &self.render_pipeline,
                &self.star_buffer,
                &self.params_buffer,
                &self.scene_rt,
            );
            if let Ok(present) = Self::bind_frame(
                &mut scheme,
                &self.scene_rt,
                self.surface.as_ref(),
                self.readback.as_ref(),
            ) {
                self.present = present;
                self.scheme = scheme;
            }
        }
    }

    fn new(window: Option<Arc<Window>>) -> Result<Self> {
        let instance = Instance::new()?;
        let device = Arc::new(
            instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let compute_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/starfield_update.slang"),
            &[Star::GPU_TYPE, StarfieldParams::GPU_TYPE],
        )?;
        let render_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/starfield_render.slang"),
            &[Star::GPU_TYPE],
        )?;

        let mut stars = Vec::with_capacity(NUM_STARS as usize);
        for _ in 0..NUM_STARS {
            let type_roll = rand_f32();
            let star_type = if type_roll < 0.70 {
                STAR_TYPE_NORMAL
            } else if type_roll < 0.85 {
                STAR_TYPE_GALAXY
            } else if type_roll < 0.90 {
                STAR_TYPE_QUASAR
            } else {
                STAR_TYPE_WHITE_DWARF
            };

            stars.push(Star {
                x: (rand_f32() - 0.5) * 0.8,
                y: (rand_f32() - 0.5) * 0.8,
                z: 0.5 + rand_f32() * 0.5,
                star_type,
            });
        }

        let star_buffer = device.acquire_buffer_with_data(&stars, BufferKind::Scattered)?;
        let params_buffer =
            device.acquire_buffer_sized::<StarfieldParams>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let compute_pipeline = ComputePipeline::new(&device, &compute_shader)?;
        let render_pipeline = Self::create_render_pipeline(&device, &render_shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_scheme(
            &mut scheme,
            &compute_pipeline,
            &render_pipeline,
            &star_buffer,
            &params_buffer,
            &scene_rt,
        );
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let params_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<StarfieldParams>(&params_buffer, 1),
        )?;

        println!("Created starfield with {NUM_STARS} stars (Scheme + Present)");

        Ok(Self {
            window,
            device,
            ctx,
            surface,
            capture,
            readback,
            present,
            scheme,
            scene_rt,
            compute_pipeline,
            render_shader,
            render_pipeline,
            star_buffer,
            params_buffer,
            upload_scheme,
            params_deposit,
            speed: 0.01,
            frame_count: 0.0,
            start_time: std::time::Instant::now(),
        })
    }

    fn render(&mut self) -> Result<()> {
        self.frame_count += 1.0;

        let params = StarfieldParams {
            speed: self.speed,
            frame: self.frame_count,
        };

        (&self.params_deposit << &params)?;
        self.upload_scheme.submit()?;

        let mut submission = self.scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        if let Some(window) = &self.window {
            window.request_redraw();
        }
        Ok(())
    }

    fn change_speed(&mut self, delta: f32) {
        self.speed = (self.speed + delta).clamp(0.001, 0.1);
        if let Some(window) = &self.window {
            window.set_title(&format!("Goldy - Starfield (speed: {:.1})", self.speed));
        }
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count as u64
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Starfield", 1024, 768))
                    .expect("Failed to create window"),
            );

            match RenderState::new(Some(window.clone())) {
                Ok(mut state) => {
                    if let Err(e) = state.render() {
                        tracing::error!("First frame error: {e}");
                    }
                    common::reveal_window(&window);
                    self.state = Some(state);
                    window.request_redraw();
                }
                Err(e) => {
                    tracing::error!("Failed to create render state: {e}");
                    event_loop.exit();
                }
            }
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if let Some(state) = &mut self.state {
                    match event.logical_key {
                        Key::Named(NamedKey::Escape) => event_loop.exit(),
                        Key::Named(NamedKey::ArrowUp) => state.change_speed(0.005),
                        Key::Named(NamedKey::ArrowDown) => state.change_speed(-0.005),
                        _ => {}
                    }
                }
            }
            WindowEvent::Resized(size) => {
                if let Some(state) = &mut self.state {
                    if size.width > 0 && size.height > 0 {
                        let Some(surface) = state.surface.as_ref() else {
                            return;
                        };
                        let (prev_w, prev_h) = surface.size();
                        if size.width == prev_w && size.height == prev_h {
                            return;
                        }
                        let _ = surface.resize(size.width, size.height);
                        let format = surface.format();
                        if let Ok(pipeline) =
                            RenderState::create_render_pipeline(&state.device, &state.render_shader, format)
                        {
                            state.render_pipeline = pipeline;
                        }
                        state.rerecord_scheme();
                    }
                }
            }
            WindowEvent::RedrawRequested => {
                if let Some(state) = &mut self.state {
                    if let Err(e) = state.render() {
                        tracing::error!("Render error: {e}");
                    }
                }
            }
            _ => {}
        }
    }
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/starfield_update.slang:

// Starfield compute shader - updates star z-positions for 3D flying effect

import goldy_exp;

static const float STAR_TYPE_NORMAL = 0.0;
static const float STAR_TYPE_GALAXY = 1.0;
static const float STAR_TYPE_QUASAR = 2.0;
static const float STAR_TYPE_WHITE_DWARF = 3.0;


static const uint NUM_STARS = 500;

float hash(float p) {
    return frac(sin(p * 127.1) * 43758.5453);
}

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Star> STARS, StarfieldParams params, ThreadId id) {
    uint idx = id.x;
    
    if (idx >= NUM_STARS) return;
    
    Star s = STARS[idx];
    
    s.z -= params.speed;
    
    if (s.z <= 0.01) {
        float seed = float(idx) + params.frame * 0.01;
        s.x = (hash(seed) - 0.5) * 2.0;
        s.y = (hash(seed + 100.0) - 0.5) * 2.0;
        s.z = 1.0;
        
        float type_roll = hash(seed + 200.0);
        if (type_roll < 0.6) {
            s.star_type = STAR_TYPE_NORMAL;
        } else if (type_roll < 0.8) {
            s.star_type = STAR_TYPE_GALAXY;
        } else if (type_roll < 0.95) {
            s.star_type = STAR_TYPE_QUASAR;
        } else {
            s.star_type = STAR_TYPE_WHITE_DWARF;
        }
    }
    
    STARS[idx] = s;
}

shaders/starfield_render.slang:

// Starfield rendering shader
// Visualizes different celestial objects using instancing

import goldy_exp;

static const float STAR_TYPE_NORMAL = 0.0;
static const float STAR_TYPE_GALAXY = 1.0;
static const float STAR_TYPE_QUASAR = 2.0;
static const float STAR_TYPE_WHITE_DWARF = 3.0;


struct VSOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
    float2 uv : TEXCOORD0;
    float star_type : TEXCOORD1;
};

static const float2 quadVerts[6] = {
    float2(-1, -1), float2( 1, -1), float2( 1,  1),
    float2(-1, -1), float2( 1,  1), float2(-1,  1)
};

[goldy_vertex]
VSOutput vs_main(Scattered<Star> STARS, VertexId vertexID, InstanceId instanceID) {
    VSOutput output;
    
    Star s = STARS[instanceID.value];
    
    float2 screenPos = float2(s.x, s.y) / max(s.z, 0.01);
    float baseSize = 0.005 + 0.02 * (1.0 - s.z);
    
    float size = baseSize;
    float3 baseColor = float3(1.0, 1.0, 1.0);
    
    float hash1 = frac(sin(float(instanceID.value) * 127.1) * 43758.5453);
    float hash2 = frac(sin(float(instanceID.value) * 269.5) * 43758.5453);
    
    if (s.star_type == STAR_TYPE_NORMAL) {
        float temp = hash1;
        if (temp < 0.15) {
            baseColor = float3(0.7, 0.85, 1.0);
        } else if (temp < 0.4) {
            baseColor = float3(1.0, 1.0, 0.95);
        } else if (temp < 0.7) {
            baseColor = float3(1.0, 0.95, 0.8);
        } else if (temp < 0.9) {
            baseColor = float3(1.0, 0.8, 0.6);
        } else {
            baseColor = float3(1.0, 0.6, 0.5);
        }
    }
    else if (s.star_type == STAR_TYPE_GALAXY) {
        size = baseSize * 2.5;
        
        float hue = hash1 * 6.0;
        int colorIdx = int(hue);
        float blend = frac(hue);
        
        float3 galaxyColors[7] = {
            float3(1.0, 0.5, 0.7),
            float3(0.7, 0.5, 1.0),
            float3(0.5, 0.7, 1.0),
            float3(0.5, 1.0, 0.9),
            float3(0.7, 1.0, 0.6),
            float3(1.0, 0.9, 0.5),
            float3(1.0, 0.5, 0.7)
        };
        
        baseColor = lerp(galaxyColors[colorIdx], galaxyColors[colorIdx + 1], blend);
    }
    else if (s.star_type == STAR_TYPE_QUASAR) {
        size = baseSize * 1.8;
        float quasarHue = hash1;
        if (quasarHue < 0.33) {
            baseColor = float3(0.4, 1.0, 1.0);
        } else if (quasarHue < 0.66) {
            baseColor = float3(0.7, 0.6, 1.0);
        } else {
            baseColor = float3(1.0, 0.7, 1.0);
        }
    }
    else if (s.star_type == STAR_TYPE_WHITE_DWARF) {
        size = baseSize * 0.7;
        baseColor = lerp(float3(0.85, 0.9, 1.0), float3(1.0, 1.0, 1.0), hash1);
    }
    
    float2 localPos = quadVerts[vertexID.value] * size;
    
    output.position = float4(screenPos + localPos, 0.0, 1.0);
    
    float brightness = 1.0 - s.z;
    output.color = float4(baseColor * brightness, 1.0);
    
    output.uv = quadVerts[vertexID.value];
    output.star_type = s.star_type;
    
    return output;
}

[goldy_fragment]
float4 fs_main(VSOutput input) : SV_Target {
    float2 uv = input.uv;
    float dist = length(uv);
    float4 color = input.color;
    
    if (input.star_type == STAR_TYPE_NORMAL) {
        float2 absUV = abs(uv);
        
        float hBar = max(0.0, 1.0 - absUV.y * 4.0) * (1.0 - smoothstep(0.0, 1.0, absUV.x));
        float vBar = max(0.0, 1.0 - absUV.x * 4.0) * (1.0 - smoothstep(0.0, 1.0, absUV.y));
        
        float cross = max(hBar, vBar);
        float core = 1.0 - smoothstep(0.0, 0.25, dist);
        
        float glow = max(cross * 0.8, core);
        color.rgb *= glow;
        color.a = glow;
    }
    else if (input.star_type == STAR_TYPE_GALAXY) {
        float2 rotUV = float2(uv.x * 0.7 + uv.y * 0.3, uv.y * 0.8 - uv.x * 0.2);
        float ellipseDist = length(rotUV * float2(1.0, 1.5));
        float glow = 1.0 - smoothstep(0.0, 0.8, ellipseDist);
        float angle = atan2(uv.y, uv.x);
        float spiral = sin(angle * 2.0 + dist * 6.0) * 0.15 + 0.85;
        color.rgb *= glow * spiral;
        color.a = glow * 0.9;
    }
    else if (input.star_type == STAR_TYPE_QUASAR) {
        float core = 1.0 - smoothstep(0.0, 0.3, dist);
        float jet = max(0.0, 1.0 - abs(uv.x) * 4.0) * (1.0 - smoothstep(0.0, 1.0, abs(uv.y)));
        float glow = max(core, jet * 0.5);
        color.rgb *= glow * 1.5;
        color.a = glow;
    }
    else if (input.star_type == STAR_TYPE_WHITE_DWARF) {
        float glow = 1.0 - smoothstep(0.0, 0.5, dist);
        glow = pow(glow, 2.0);
        color.rgb *= glow * 1.3;
        color.a = glow;
    }
    
    if (color.a < 0.01) discard;
    
    return color;
}

particles

A rain and snow particle system. The compute pass integrates positions and wraps particles at the window edges; pressing Space swaps the simulation mode at runtime.

cargo run --features examples --example particles

What it demonstrates

  • Runtime switching between simulation modes
  • Compute dispatch feeding an instanced raster pass

Controls

KeyAction
SpaceToggle between rain and snow
EscapeExit

Source

examples/particles.rs:

//! Particles example - rain/snow particle system.
//!
//! Demonstrates retained scheme with compute dispatch → offscreen render → copy-to-present.
//!
//! Run with: `cargo run --example particles`

use anyhow::Result;
use goldy::{
    Buffer, BufferFlags, BufferKind, Color, ComputePipeline, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, PrimitiveTopology, RenderPipeline, RenderPipelineDesc,
    RequestAdapterOptions, RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad,
    Texture, TextureFormat, Transaction, VertexBufferLayout,
};
use std::ops::Shr;
use std::sync::Arc;
use winit::{
    application::ApplicationHandler,
    event::WindowEvent,
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowId},
};
mod common;
use common::CaptureDump;

const NUM_PARTICLES: u32 = 1000;

#[goldy::gpu]
struct Particle {
    position: [f32; 2],
    velocity: [f32; 2],
    size: f32,
}

#[goldy::gpu]
struct ParticleParams {
    is_snow: f32,
    frame: f32,
}

static mut SEED: u32 = 42;
fn random() -> f32 {
    unsafe {
        SEED = SEED.wrapping_mul(1103515245).wrapping_add(12345);
        SEED as f32 / u32::MAX as f32
    }
}

fn main() -> Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        let mut state = RenderState::new(None)?;
        while !state.capture_done() {
            state.render()?;
        }
        return Ok(());
    }

    println!("Goldy Particles Example");
    println!("  Space - Toggle rain/snow");
    println!("  Escape - Exit");

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);

    let mut app = App::default();
    event_loop.run_app(&mut app)?;

    Ok(())
}

#[derive(Default)]
struct App {
    state: Option<RenderState>,
}

struct RenderState {
    window: Option<Arc<Window>>,
    device: Arc<goldy::Runtime>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    present: Option<Transaction>,
    scheme: Scheme,
    scene_rt: Lease<LeaseRenderTarget>,
    compute_pipeline: ComputePipeline,
    render_shader: ShaderModule,
    render_pipeline: RenderPipeline,
    particle_buffer: Buffer,
    params_buffer: Buffer,
    upload_scheme: Scheme,
    params_deposit: DepositTransaction,
    is_snow: bool,
    frame_count: f32,
    start_time: std::time::Instant,
}

impl RenderState {
    fn create_render_pipeline(
        device: &goldy::Runtime,
        render_shader: &ShaderModule,
        format: TextureFormat,
    ) -> Result<RenderPipeline> {
        common::render_pipeline(
            device,
            render_shader,
            format,
            RenderPipelineDesc {
                vertex_layout: VertexBufferLayout::empty(),
                topology: PrimitiveTopology::TriangleList,
                ..Default::default()
            },
        )
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn record_scheme(
        scheme: &mut Scheme,
        compute_pipeline: &ComputePipeline,
        render_pipeline: &RenderPipeline,
        particle_buffer: &Buffer,
        params_buffer: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
        bg_color: Color,
    ) {
        scheme
            .node("update_particles", compute_pipeline)
            .with_parcel(particle_buffer, NodeAccess::ReadWrite)
            .with_parcel(params_buffer, NodeAccess::Read)
            .dispatch(NUM_PARTICLES.div_ceil(64), 1, 1);

        let mut pass = scheme.render_pass("particles", scene_rt, TargetLoad::Clear(bg_color));
        pass.with_parcel(particle_buffer, NodeAccess::Read);
        pass.with_parcel(params_buffer, NodeAccess::Read);
        pass.set_pipeline(render_pipeline);
        pass.draw(0..6, 0..NUM_PARTICLES);
        pass.finish();
    }

    fn background_color(is_snow: bool) -> Color {
        if is_snow {
            Color {
                r: 0.05,
                g: 0.05,
                b: 0.15,
                a: 1.0,
            }
        } else {
            Color {
                r: 0.02,
                g: 0.02,
                b: 0.05,
                a: 1.0,
            }
        }
    }

    fn target(&self) -> (TextureFormat, u32, u32) {
        if let Some(surface) = &self.surface {
            let (width, height) = surface.size();
            (surface.format(), width, height)
        } else {
            let capture = self.capture.as_ref().expect("capture dump");
            let (width, height) = capture.size();
            (CaptureDump::format(), width, height)
        }
    }

    fn capture_done(&self) -> bool {
        self.capture.as_ref().is_none_or(CaptureDump::finished)
    }

    fn rerecord_scheme(&mut self) {
        let mut scheme = Scheme::new(&self.ctx);
        let (format, width, height) = self.target();
        if let Ok(rt) = self.ctx.lease_render_target(width.max(1), height.max(1), format, None) {
            self.scene_rt = rt;
            Self::record_scheme(
                &mut scheme,
                &self.compute_pipeline,
                &self.render_pipeline,
                &self.particle_buffer,
                &self.params_buffer,
                &self.scene_rt,
                Self::background_color(self.is_snow),
            );
            if let Ok(present) = Self::bind_frame(
                &mut scheme,
                &self.scene_rt,
                self.surface.as_ref(),
                self.readback.as_ref(),
            ) {
                self.present = present;
                self.scheme = scheme;
            }
        }
    }

    fn new(window: Option<Arc<Window>>) -> Result<Self> {
        let instance = Instance::new()?;
        let device = Arc::new(
            instance
                .request_adapter(&RequestAdapterOptions::default())?
                .request_runtime(&RuntimeDescriptor::default())?,
        );
        let ctx = device.create_context()?;

        let (surface, capture, readback, format, width, height) = if let Some(window) = window.as_deref() {
            let surface = SurfaceExchange::new(&ctx, window, SurfaceConfig::default())?;
            let format = surface.format();
            let (width, height) = surface.size();
            (Some(surface), None, None, format, width, height)
        } else {
            let capture = CaptureDump::from_env()?;
            let (width, height) = capture.size();
            let readback = common::capture_readback(&device, width, height)?;
            (
                None,
                Some(capture),
                Some(readback),
                CaptureDump::format(),
                width,
                height,
            )
        };

        let compute_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/rain_snow_update.slang"),
            &[Particle::GPU_TYPE, ParticleParams::GPU_TYPE],
        )?;
        let render_shader = ShaderModule::from_slang_with_gpu_types(
            &device,
            include_str!("../shaders/rain_snow_render.slang"),
            &[Particle::GPU_TYPE, ParticleParams::GPU_TYPE],
        )?;

        let particles = Self::create_particles(false);
        let particle_buffer = device.acquire_buffer_with_data(&particles, BufferKind::Scattered)?;
        let params_buffer =
            device.acquire_buffer_sized::<ParticleParams>(1, BufferKind::Broadcast, BufferFlags::empty())?;

        let compute_pipeline = ComputePipeline::new(&device, &compute_shader)?;
        let render_pipeline = Self::create_render_pipeline(&device, &render_shader, format)?;

        let mut scheme = Scheme::new(&ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_scheme(
            &mut scheme,
            &compute_pipeline,
            &render_pipeline,
            &particle_buffer,
            &params_buffer,
            &scene_rt,
            Self::background_color(false),
        );
        let present = Self::bind_frame(&mut scheme, &scene_rt, surface.as_ref(), readback.as_ref())?;

        let mut upload_scheme = Scheme::new(&ctx);
        let params_deposit = MemoryExchange::new(&ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer_elements::<ParticleParams>(&params_buffer, 1),
        )?;

        println!("Created rain/snow simulation with {NUM_PARTICLES} particles (Scheme + Present)");

        Ok(Self {
            window,
            device,
            ctx,
            surface,
            capture,
            readback,
            present,
            scheme,
            scene_rt,
            compute_pipeline,
            render_shader,
            render_pipeline,
            particle_buffer,
            params_buffer,
            upload_scheme,
            params_deposit,
            is_snow: false,
            frame_count: 0.0,
            start_time: std::time::Instant::now(),
        })
    }

    fn create_particles(is_snow: bool) -> Vec<Particle> {
        let mut particles = Vec::with_capacity(NUM_PARTICLES as usize);
        for _ in 0..NUM_PARTICLES {
            let x = random() * 2.0 - 1.0;
            let y = random() * 2.2 - 1.0;

            let (vx, vy, size) = if is_snow {
                (
                    (random() - 0.5) * 0.005,
                    -(0.002 + random() * 0.005),
                    0.003 + random() * 0.008,
                )
            } else {
                (
                    (random() - 0.5) * 0.002,
                    -(0.01 + random() * 0.02),
                    0.002 + random() * 0.003,
                )
            };

            particles.push(Particle {
                position: [x, y],
                velocity: [vx, vy],
                size,
            });
        }
        particles
    }

    fn toggle_mode(&mut self) -> Result<()> {
        self.is_snow = !self.is_snow;

        let particles = Self::create_particles(self.is_snow);
        let mut particle_upload = Scheme::new(&self.ctx);
        let particle_deposit = MemoryExchange::new(&self.ctx).bind_deposit(
            &mut particle_upload,
            DepositTarget::buffer_elements::<Particle>(&self.particle_buffer, NUM_PARTICLES as u64),
        )?;
        (&particle_deposit << particles.as_slice())?;
        particle_upload.submit()?;

        if let Some(window) = &self.window {
            window.set_title(&format!(
                "Goldy - {} (Space to toggle)",
                if self.is_snow { "Snow" } else { "Rain" }
            ));
        }

        self.rerecord_scheme();
        Ok(())
    }

    fn render(&mut self) -> Result<()> {
        self.frame_count += 1.0;

        let params = ParticleParams {
            is_snow: if self.is_snow { 1.0 } else { 0.0 },
            frame: self.frame_count,
        };

        (&self.params_deposit << &params)?;
        self.upload_scheme.submit()?;

        let mut submission = self.scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }

        if let Some(window) = &self.window {
            window.request_redraw();
        }
        Ok(())
    }
}

impl Drop for RenderState {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count as u64
        );
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.state.is_none() {
            let window = Arc::new(
                event_loop
                    .create_window(common::hidden_window("Goldy - Rain (Space to toggle)", 800, 600))
                    .expect("Failed to create window"),
            );

            match RenderState::new(Some(window.clone())) {
                Ok(mut state) => {
                    if let Err(e) = state.render() {
                        tracing::error!("First frame error: {e}");
                    }
                    common::reveal_window(&window);
                    self.state = Some(state);
                    window.request_redraw();
                }
                Err(e) => {
                    tracing::error!("Failed to create render state: {e:#}");
                    event_loop.exit();
                }
            }
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        if let Some(state) = &self.state {
            common::exit_if_timed_out(event_loop, state.start_time);
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, _id: WindowId, event: WindowEvent) {
        match event {
            WindowEvent::CloseRequested => event_loop.exit(),
            WindowEvent::KeyboardInput { event, .. } if event.state.is_pressed() => {
                if let Some(state) = &mut self.state {
                    match event.logical_key {
                        Key::Named(NamedKey::Escape) => event_loop.exit(),
                        Key::Named(NamedKey::Space) => {
                            if let Err(e) = state.toggle_mode() {
                                tracing::error!("Failed to toggle mode: {e}");
                            }
                        }
                        _ => {}
                    }
                }
            }
            WindowEvent::Resized(size) => {
                if let Some(state) = &mut self.state {
                    if size.width > 0 && size.height > 0 {
                        let Some(surface) = state.surface.as_ref() else {
                            return;
                        };
                        let (prev_w, prev_h) = surface.size();
                        if size.width == prev_w && size.height == prev_h {
                            return;
                        }
                        let _ = surface.resize(size.width, size.height);
                        let format = surface.format();
                        if let Ok(pipeline) =
                            RenderState::create_render_pipeline(&state.device, &state.render_shader, format)
                        {
                            state.render_pipeline = pipeline;
                        }
                        state.rerecord_scheme();
                    }
                }
            }
            WindowEvent::RedrawRequested => {
                if let Some(state) = &mut self.state {
                    if let Err(e) = state.render() {
                        tracing::error!("Render error: {e}");
                    }
                }
            }
            _ => {}
        }
    }
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/rain_snow_update.slang:

// Rain/Snow particle compute shader
// Updates particle positions with mode-dependent physics

import goldy_exp;


static const uint NUM_PARTICLES = 1000;

float hash(float p) {
    return frac(sin(p * 127.1) * 43758.5453);
}

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Particle> PARTICLES, ParticleParams params, ThreadId id) {
    uint idx = id.x;
    
    if (idx >= NUM_PARTICLES) return;
    
    Particle p = PARTICLES[idx];
    bool is_snow = params.is_snow > 0.5;
    
    p.position += p.velocity;
    
    if (is_snow) {
        float seed = float(idx) + params.frame * 0.1;
        p.velocity.x += (hash(seed) - 0.5) * 0.001;
        p.velocity.x = clamp(p.velocity.x, -0.01, 0.01);
    }
    
    if (p.position.y < -1.2 || p.position.x < -1.2 || p.position.x > 1.2) {
        float seed = float(idx) + params.frame * 0.01;
        
        p.position.x = hash(seed) * 2.0 - 1.0;
        p.position.y = 1.0 + hash(seed + 100.0) * 0.5;
        
        if (is_snow) {
            p.velocity.x = (hash(seed + 200.0) - 0.5) * 0.005;
            p.velocity.y = -(0.002 + hash(seed + 300.0) * 0.005);
            p.size = 0.003 + hash(seed + 400.0) * 0.008;
        } else {
            p.velocity.x = (hash(seed + 200.0) - 0.5) * 0.002;
            p.velocity.y = -(0.01 + hash(seed + 300.0) * 0.02);
            p.size = 0.002 + hash(seed + 400.0) * 0.003;
        }
    }
    
    PARTICLES[idx] = p;
}

shaders/rain_snow_render.slang:

// Rain/Snow particle rendering shader
// Visualizes particles as elongated quads using instancing

import goldy_exp;


struct VSOutput {
    float4 position : SV_Position;
    float4 color : COLOR;
};

static const float2 quadVerts[6] = {
    float2(-1, -3), float2( 1, -3), float2( 1,  3),
    float2(-1, -3), float2( 1,  3), float2(-1,  3)
};

static const float2 snowQuadVerts[6] = {
    float2(-1, -1), float2( 1, -1), float2( 1,  1),
    float2(-1, -1), float2( 1,  1), float2(-1,  1)
};

[goldy_vertex]
VSOutput vs_main(Scattered<Particle> PARTICLES, ParticleParams params,
                 VertexId vertexID, InstanceId instanceID) {
    VSOutput output;
    
    Particle p = PARTICLES[instanceID.value];
    bool is_snow = params.is_snow > 0.5;
    
    float size = p.size;
    float2 localPos;
    
    if (is_snow) {
        localPos = snowQuadVerts[vertexID.value] * size;
    } else {
        localPos = quadVerts[vertexID.value] * size;
    }
    
    output.position = float4(p.position + localPos, 0.0, 1.0);
    
    if (is_snow) {
        float brightness = 0.7 + size * 30.0;
        output.color = float4(0.95 * brightness, 0.95 * brightness, 1.0 * brightness, 0.8);
    } else {
        float brightness = 0.6 + size * 50.0;
        output.color = float4(0.5 * brightness, 0.6 * brightness, 0.9 * brightness, 0.7);
    }
    
    return output;
}

[goldy_fragment]
float4 fs_main(VSOutput input) : SV_Target {
    return input.color;
}

multi_window

Three windows — plasma, tunnel, and starfield — sharing one Runtime. Each window owns its own SurfaceExchange, Context, and Scheme, which is the pattern for any multi-surface application.

cargo run --features examples --example multi_window

What it demonstrates

  • Multiple surfaces on a single device
  • One scheme and context per window
  • Independent resize and close handling per window

Controls

KeyAction
SpaceToggle the focused window's effect modifier
RReset the focused window's effect
EscapeClose the focused window; the app exits with the last one

Source

examples/multi_window.rs:

//! Multi-window example - three simultaneous effects in separate windows.
//!
//! Each window runs its own demo with an independent SurfaceExchange + Scheme.
//!
//! Run with: cargo run --example multi_window

use goldy::{
    shaders, Buffer, BufferFlags, BufferKind, Color, DepositTarget, DepositTransaction, Instance, Lease,
    LeaseRenderTarget, MemoryExchange, NodeAccess, RenderPipeline, RenderPipelineDesc, RequestAdapterOptions,
    RuntimeDescriptor, Scheme, ShaderModule, SurfaceConfig, SurfaceExchange, TargetLoad, Texture, TextureFormat,
    Transaction, VertexBufferLayout,
};
mod common;
use common::CaptureDump;
use std::ops::Shr;

const PLASMA_VERTEX_TIME: &str = r#"
struct VertexInput {
    float2 position : POSITION;
    float2 uv : TEXCOORD0;
    float time : TEXCOORD1;
};

struct VertexOutput {
    float4 position : SV_Position;
    float2 uv : TEXCOORD0;
    float time : TEXCOORD1;
};

[shader("vertex")]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.uv = input.uv;
    output.time = input.time;
    return output;
}

float3 rainbow(float t) {
    float3 c = float3(
        sin(t * 6.28318 + 0.0) * 0.5 + 0.5,
        sin(t * 6.28318 + 2.094) * 0.5 + 0.5,
        sin(t * 6.28318 + 4.189) * 0.5 + 0.5
    );
    return c;
}

[shader("fragment")]
float4 fs_main(VertexOutput input) : SV_Target {
    float2 uv = input.uv * 4.0;
    float t = input.time;
    
    float v = sin(uv.x + t);
    v += sin(uv.y + t);
    v += sin(uv.x + uv.y + t);
    
    float cx = uv.x + 0.5 * sin(t / 3.0);
    float cy = uv.y + 0.5 * cos(t / 2.0);
    v += sin(sqrt(cx * cx + cy * cy + 1.0) + t);
    
    v = v / 2.0;
    
    return float4(rainbow(v), 1.0);
}
"#;

const TUNNEL_VERTEX_TIME: &str = r#"
struct VertexInput {
    float2 position : POSITION;
    float2 uv : TEXCOORD0;
    float time : TEXCOORD1;
};

struct VertexOutput {
    float4 position : SV_Position;
    float2 uv : TEXCOORD0;
    float time : TEXCOORD1;
};

[shader("vertex")]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.uv = input.uv;
    output.time = input.time;
    return output;
}

[shader("fragment")]
float4 fs_main(VertexOutput input) : SV_Target {
    float2 uv = (input.uv - 0.5) * 2.0;
    float t = input.time;
    
    float dist = length(uv);
    float angle = atan2(uv.y, uv.x);
    
    float tunnel_depth = 1.0 / (dist + 0.1);
    float tunnel_angle = angle / 3.14159 + t * 0.2;
    
    float tx = tunnel_angle * 4.0;
    float ty = tunnel_depth - t * 2.0;
    
    float checker = floor(tx) + floor(ty);
    bool is_white = fmod(checker, 2.0) == 0.0;
    
    float depth_color = 1.0 - dist * 0.5;
    float3 color;
    
    if (is_white) {
        color = float3(0.8, 0.2, 0.4) * depth_color;
    } else {
        color = float3(0.2, 0.4, 0.8) * depth_color;
    }
    
    color += float3(0.3, 0.5, 1.0) * (1.0 - dist) * (1.0 - dist);
    color *= 1.0 - dist * 0.3;
    
    return float4(color, 1.0);
}
"#;

use std::collections::HashMap;
use std::sync::Arc;
use std::time::Instant;
use winit::{
    application::ApplicationHandler,
    dpi::LogicalSize,
    event::{ElementState, MouseButton, WindowEvent},
    event_loop::{ActiveEventLoop, ControlFlow, EventLoop},
    keyboard::{Key, NamedKey},
    window::{Window, WindowAttributes, WindowId},
};
#[goldy::gpu]
struct QuadVertex {
    position: [f32; 2],
    uv: [f32; 2],
    time: f32,
}

impl QuadVertex {
    fn layout() -> VertexBufferLayout {
        Self::GPU_TYPE.vertex_buffer_layout().expect("quad vertex layout")
    }
}

fn create_quad(time: f32) -> [QuadVertex; 6] {
    [
        QuadVertex {
            position: [-1.0, -1.0],
            uv: [0.0, 1.0],
            time,
        },
        QuadVertex {
            position: [1.0, -1.0],
            uv: [1.0, 1.0],
            time,
        },
        QuadVertex {
            position: [1.0, 1.0],
            uv: [1.0, 0.0],
            time,
        },
        QuadVertex {
            position: [-1.0, -1.0],
            uv: [0.0, 1.0],
            time,
        },
        QuadVertex {
            position: [1.0, 1.0],
            uv: [1.0, 0.0],
            time,
        },
        QuadVertex {
            position: [-1.0, 1.0],
            uv: [0.0, 0.0],
            time,
        },
    ]
}

#[derive(Clone, Copy, PartialEq)]
enum EffectType {
    Plasma,
    Tunnel,
    Starfield,
}

impl EffectType {
    fn title(&self) -> &'static str {
        match self {
            EffectType::Plasma => "Plasma [Space=pause, Click=reset]",
            EffectType::Tunnel => "Tunnel [Space=reverse, Click=reset]",
            EffectType::Starfield => "Starfield [Space=warp, Click=reset]",
        }
    }

    fn shader_source(&self) -> &'static str {
        match self {
            EffectType::Plasma => PLASMA_VERTEX_TIME,
            EffectType::Tunnel => TUNNEL_VERTEX_TIME,
            EffectType::Starfield => shaders::STARFIELD,
        }
    }
}

struct WindowState {
    window: Option<Arc<Window>>,
    ctx: goldy::Context,
    surface: Option<SurfaceExchange>,
    present: Option<Transaction>,
    capture: Option<CaptureDump>,
    readback: Option<Texture>,
    scheme: Scheme,
    scene_rt: Lease<LeaseRenderTarget>,
    pipeline: RenderPipeline,
    effect_type: EffectType,
    start_time: Instant,
    paused: bool,
    paused_at: f32,
    time_multiplier: f32,
    vertex_parcel: Buffer,
    upload_scheme: Scheme,
    vertex_deposit: DepositTransaction,
    has_focus: bool,
}

impl WindowState {
    fn create_pipeline(
        device: &goldy::Runtime,
        shader: &ShaderModule,
        format: TextureFormat,
    ) -> anyhow::Result<RenderPipeline> {
        common::render_pipeline(
            device,
            shader,
            format,
            RenderPipelineDesc {
                vertex_layout: QuadVertex::layout(),
                ..Default::default()
            },
        )
    }

    fn record_pass(
        scheme: &mut Scheme,
        pipeline: &RenderPipeline,
        vertex_parcel: &Buffer,
        scene_rt: &Lease<LeaseRenderTarget>,
        label: &'static str,
    ) {
        let mut pass = scheme.render_pass(label, scene_rt, TargetLoad::Clear(Color::BLACK));
        pass.with_parcel(vertex_parcel, NodeAccess::Read);
        pass.set_pipeline(pipeline);
        pass.set_vertex_buffer(0, vertex_parcel);
        pass.draw(0..6, 0..1);
        pass.finish();
    }

    fn bind_frame(
        scheme: &mut Scheme,
        scene_rt: &Lease<LeaseRenderTarget>,
        surface: Option<&SurfaceExchange>,
        readback: Option<&Texture>,
    ) -> anyhow::Result<Option<Transaction>> {
        if let Some(surface) = surface {
            let present = surface.bind_render_target(scheme, scene_rt)?;
            Ok(Some(present))
        } else {
            let readback = readback.expect("capture readback");
            scheme.copy_to_texture(scene_rt, readback)?;

            Ok(None)
        }
    }

    fn output_size(&self) -> (u32, u32) {
        if let Some(surface) = &self.surface {
            surface.size()
        } else {
            self.capture.as_ref().expect("capture").size()
        }
    }

    fn output_format(&self) -> TextureFormat {
        if let Some(surface) = &self.surface {
            surface.format()
        } else {
            CaptureDump::format()
        }
    }

    fn rerecord_scheme(&mut self) {
        let mut scheme = Scheme::new(&self.ctx);
        let (width, height) = self.output_size();
        let format = self.output_format();
        if let Ok(rt) = self.ctx.lease_render_target(width.max(1), height.max(1), format, None) {
            Self::record_pass(
                &mut scheme,
                &self.pipeline,
                &self.vertex_parcel,
                &rt,
                self.effect_type.title(),
            );
            if let Ok(present) = Self::bind_frame(&mut scheme, &rt, self.surface.as_ref(), self.readback.as_ref()) {
                self.present = present;
                self.scene_rt = rt;
                self.scheme = scheme;
            }
        }
    }

    fn windowed(
        window: Arc<Window>,
        ctx: &goldy::Context,
        device: &Arc<goldy::Runtime>,
        effect_type: EffectType,
    ) -> anyhow::Result<Self> {
        let surface = SurfaceExchange::new(ctx, window.as_ref(), SurfaceConfig::default())?;
        let format = surface.format();
        let (width, height) = surface.size();
        let shader = ShaderModule::from_slang(device, effect_type.shader_source())?;
        let pipeline = Self::create_pipeline(device, &shader, format)?;

        let vertex_parcel =
            device.acquire_buffer_sized::<QuadVertex>(6, BufferKind::Scattered, BufferFlags::empty())?;

        let mut scheme = Scheme::new(ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &vertex_parcel, &scene_rt, effect_type.title());
        let present = Self::bind_frame(&mut scheme, &scene_rt, Some(&surface), None)?;

        let mut upload_scheme = Scheme::new(ctx);
        let vertex_deposit = MemoryExchange::new(ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer(&vertex_parcel, vertex_parcel.byte_size()),
        )?;

        Ok(Self {
            window: Some(window),
            ctx: ctx.clone(),
            surface: Some(surface),
            present,
            capture: None,
            readback: None,
            scheme,
            scene_rt,
            pipeline,
            effect_type,
            start_time: Instant::now(),
            paused: false,
            paused_at: 0.0,
            time_multiplier: 1.0,
            vertex_parcel,
            upload_scheme,
            vertex_deposit,
            has_focus: false,
        })
    }

    fn capture_panel(
        ctx: &goldy::Context,
        device: &Arc<goldy::Runtime>,
        effect_type: EffectType,
        width: u32,
        height: u32,
    ) -> anyhow::Result<Self> {
        let capture = CaptureDump::memory(width, height);
        let format = CaptureDump::format();
        let readback = common::capture_readback(&device, width, height)?;
        let shader = ShaderModule::from_slang(device, effect_type.shader_source())?;
        let pipeline = Self::create_pipeline(device, &shader, format)?;

        let vertex_parcel =
            device.acquire_buffer_sized::<QuadVertex>(6, BufferKind::Scattered, BufferFlags::empty())?;

        let mut scheme = Scheme::new(ctx);
        let scene_rt = ctx.lease_render_target(width.max(1), height.max(1), format, None)?;
        Self::record_pass(&mut scheme, &pipeline, &vertex_parcel, &scene_rt, effect_type.title());
        let present = Self::bind_frame(&mut scheme, &scene_rt, None, Some(&readback))?;

        let mut upload_scheme = Scheme::new(ctx);
        let vertex_deposit = MemoryExchange::new(ctx).bind_deposit(
            &mut upload_scheme,
            DepositTarget::buffer(&vertex_parcel, vertex_parcel.byte_size()),
        )?;

        Ok(Self {
            window: None,
            ctx: ctx.clone(),
            surface: None,
            present,
            capture: Some(capture),
            readback: Some(readback),
            scheme,
            scene_rt,
            pipeline,
            effect_type,
            start_time: Instant::now(),
            paused: false,
            paused_at: 0.0,
            time_multiplier: 1.0,
            vertex_parcel,
            upload_scheme,
            vertex_deposit,
            has_focus: false,
        })
    }

    fn current_time(&self) -> f32 {
        if let Some(capture) = &self.capture {
            return capture.time();
        }
        if self.paused {
            self.paused_at
        } else {
            self.paused_at + self.start_time.elapsed().as_secs_f32() * self.time_multiplier
        }
    }

    fn toggle_pause(&mut self) {
        if self.paused {
            self.start_time = Instant::now();
            self.paused = false;
        } else {
            self.paused_at = self.current_time();
            self.paused = true;
        }
    }

    fn toggle_effect_modifier(&mut self) {
        match self.effect_type {
            EffectType::Plasma => self.toggle_pause(),
            EffectType::Tunnel => {
                self.paused_at = self.current_time();
                self.start_time = Instant::now();
                self.time_multiplier *= -1.0;
            }
            EffectType::Starfield => {
                self.paused_at = self.current_time();
                self.start_time = Instant::now();
                self.time_multiplier = if self.time_multiplier > 2.0 { 1.0 } else { 5.0 };
            }
        }
    }

    fn reset(&mut self) {
        self.start_time = Instant::now();
        self.paused = false;
        self.paused_at = 0.0;
        self.time_multiplier = 1.0;
    }

    fn render(&mut self, _ctx: &goldy::Context) -> anyhow::Result<()> {
        if let Some(window) = &self.window {
            let size = window.inner_size();
            if size.width == 0 || size.height == 0 {
                return Ok(());
            }
        }

        let vertices = create_quad(self.current_time());
        (&self.vertex_deposit << vertices.as_slice())?;
        self.upload_scheme.submit()?;

        let mut submission = self.scheme.submit()?;
        if let Some(present) = &self.present {
            (&mut submission >> present).take()?;
        } else {
            let pixels = (&mut submission >> self.readback.as_ref().unwrap())
                .take::<u8>()?
                .to_vec();
            self.capture.as_mut().unwrap().write_rgba(&pixels)?;
        }
        Ok(())
    }

    fn handle_resize(&mut self, width: u32, height: u32) {
        if width == 0 || height == 0 {
            return;
        }
        let Some(surface) = &self.surface else {
            return;
        };
        let (prev_w, prev_h) = surface.size();
        // Pipeline does not depend on surface size; skip no-op Resized events
        // (winit often fires these on reveal) so we don't recompile Slang→DXIL.
        if prev_w == width && prev_h == height {
            return;
        }
        let _ = surface.resize(width, height);
        self.rerecord_scheme();
    }
}

struct App {
    instance: Instance,
    ctx: Option<goldy::Context>,
    device: Option<Arc<goldy::Runtime>>,
    windows: HashMap<WindowId, WindowState>,
    effects_to_create: Vec<EffectType>,
    frame_count: u32,
    start_time: std::time::Instant,
}

impl App {
    fn new() -> anyhow::Result<Self> {
        Ok(Self {
            instance: Instance::new()?,
            ctx: None,
            device: None,
            windows: HashMap::new(),
            effects_to_create: vec![EffectType::Plasma, EffectType::Tunnel, EffectType::Starfield],
            frame_count: 0,
            start_time: std::time::Instant::now(),
        })
    }

    fn create_window(
        &mut self,
        event_loop: &ActiveEventLoop,
        effect_type: EffectType,
        position: (i32, i32),
    ) -> anyhow::Result<()> {
        let device = self.device.as_ref().unwrap().clone();
        let ctx = self.ctx.as_ref().unwrap();

        let attrs = WindowAttributes::default()
            .with_title(format!("Goldy - {}", effect_type.title()))
            .with_inner_size(LogicalSize::new(500, 500))
            .with_position(winit::dpi::LogicalPosition::new(position.0, position.1))
            .with_visible(false);

        let window = Arc::new(event_loop.create_window(attrs)?);
        let window_id = window.id();

        let mut state = WindowState::windowed(window.clone(), ctx, &device, effect_type)?;
        state.render(ctx)?;
        common::reveal_window(&window);
        window.request_redraw();

        self.windows.insert(window_id, state);
        Ok(())
    }
}

impl ApplicationHandler for App {
    fn resumed(&mut self, event_loop: &ActiveEventLoop) {
        if self.device.is_none() {
            match self
                .instance
                .request_adapter(&RequestAdapterOptions::default())
                .and_then(|a| a.request_runtime(&RuntimeDescriptor::default()))
            {
                Ok(device) => {
                    let device = Arc::new(device);
                    self.ctx = Some(device.create_context().expect("create context"));
                    self.device = Some(device);
                }
                Err(e) => {
                    tracing::error!("Failed to create device: {}", e);
                    event_loop.exit();
                    return;
                }
            }
        }

        let effects = std::mem::take(&mut self.effects_to_create);
        for (i, effect) in effects.into_iter().enumerate() {
            let x = 50 + (i as i32) * 520;
            let y = 100;

            if let Err(e) = self.create_window(event_loop, effect, (x, y)) {
                tracing::error!("Failed to create window for {:?}: {}", effect.title(), e);
            }
        }
    }

    fn window_event(&mut self, event_loop: &ActiveEventLoop, window_id: WindowId, event: WindowEvent) {
        let state = match self.windows.get_mut(&window_id) {
            Some(s) => s,
            None => return,
        };

        match event {
            WindowEvent::CloseRequested => {
                self.windows.remove(&window_id);
                if self.windows.is_empty() {
                    event_loop.exit();
                }
            }
            WindowEvent::Focused(focused) => {
                state.has_focus = focused;
                if focused {
                    println!(
                        "Focus: {} ({})",
                        state.effect_type.title(),
                        if state.paused { "paused" } else { "running" }
                    );
                }
            }
            WindowEvent::KeyboardInput { event, .. } if event.state == ElementState::Pressed => {
                match event.logical_key {
                    Key::Named(NamedKey::Escape) => {
                        self.windows.remove(&window_id);
                        if self.windows.is_empty() {
                            event_loop.exit();
                        }
                    }
                    Key::Named(NamedKey::Space) => {
                        if let Some(s) = self.windows.get_mut(&window_id) {
                            s.toggle_effect_modifier();
                            println!("[{}] Modifier toggled", s.effect_type.title());
                        }
                    }
                    Key::Character(ref c) if c == "r" || c == "R" => {
                        if let Some(s) = self.windows.get_mut(&window_id) {
                            s.reset();
                            println!("[{}] Reset", s.effect_type.title());
                        }
                    }
                    _ => {}
                }
            }
            WindowEvent::MouseInput {
                state: ElementState::Pressed,
                button: MouseButton::Left,
                ..
            } => {
                if let Some(s) = self.windows.get_mut(&window_id) {
                    s.reset();
                    println!("[{}] Reset (click)", s.effect_type.title());
                }
            }
            WindowEvent::RedrawRequested => {
                let Some(ctx) = self.ctx.clone() else {
                    return;
                };
                if let Some(s) = self.windows.get_mut(&window_id) {
                    if let Err(e) = s.render(&ctx) {
                        tracing::error!("[{}] Render error: {}", s.effect_type.title(), e);
                    }
                }
            }
            WindowEvent::Resized(new_size) => {
                if let Some(s) = self.windows.get_mut(&window_id) {
                    s.handle_resize(new_size.width, new_size.height);
                }
            }
            _ => {}
        }
    }

    fn about_to_wait(&mut self, event_loop: &ActiveEventLoop) {
        common::exit_if_timed_out(event_loop, self.start_time);

        if self.ctx.is_none() {
            return;
        }

        self.frame_count += 1;

        for state in self.windows.values() {
            if let Some(window) = &state.window {
                window.request_redraw();
            }
        }
    }
}

impl Drop for App {
    fn drop(&mut self) {
        let elapsed = self.start_time.elapsed().as_secs_f64();
        let fps = if elapsed > 0.0 {
            self.frame_count as f64 / elapsed
        } else {
            0.0
        };
        println!(
            "GOLDY_PERF: frames={} elapsed={elapsed:.2}s avg_fps={fps:.1}",
            self.frame_count
        );
    }
}

fn capture_panels() -> anyhow::Result<()> {
    let instance = Instance::new()?;
    let device = Arc::new(
        instance
            .request_adapter(&RequestAdapterOptions::default())?
            .request_runtime(&RuntimeDescriptor::default())?,
    );
    let ctx = device.create_context()?;
    let mut output = CaptureDump::from_env()?;
    let (out_w, out_h) = output.size();
    anyhow::ensure!(out_w % 3 == 0, "multi_window capture width must be divisible by 3");
    let panel = (out_w / 3, out_h);
    let effects = [EffectType::Plasma, EffectType::Tunnel, EffectType::Starfield];
    let mut panels = Vec::new();
    for effect in effects {
        panels.push(WindowState::capture_panel(&ctx, &device, effect, panel.0, panel.1)?);
    }
    while !output.finished() {
        let mut frames = Vec::with_capacity(3);
        for panel_state in &mut panels {
            panel_state.render(&ctx)?;
            frames.push(
                panel_state
                    .capture
                    .as_mut()
                    .and_then(CaptureDump::take_rgba)
                    .ok_or_else(|| anyhow::anyhow!("missing panel pixels"))?,
            );
        }
        let stacked = common::hstack_rgba(&[&frames[0], &frames[1], &frames[2]], panel.0, panel.1)?;
        output.write_rgba(&stacked)?;
    }
    Ok(())
}

fn main() -> anyhow::Result<()> {
    tracing_subscriber::fmt()
        .with_env_filter(
            tracing_subscriber::EnvFilter::try_from_default_env()
                .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("warn")),
        )
        .init();

    if common::capture_requested() {
        return capture_panels();
    }

    println!("Goldy Multi-Window Example (Scheme + Present)");
    println!("Three windows, three effects, independent controls:");
    println!();
    println!("  Plasma:    Space=pause     Click/R=reset");
    println!("  Tunnel:    Space=reverse   Click/R=reset");
    println!("  Starfield: Space=warp      Click/R=reset");
    println!();
    println!("Escape closes the focused window. Close all to exit.");
    println!();

    let event_loop = EventLoop::new()?;
    event_loop.set_control_flow(ControlFlow::Poll);
    event_loop.run_app(&mut App::new()?)?;
    Ok(())
}

The example pulls in examples/common.rs — see Shared Helpers.

Shaders

shaders/starfield.slang:

// 3D Starfield - flying forward through space
// Uses vertex time attribute for animation (same pattern as plasma/tunnel)

struct VertexInput {
    float2 position : POSITION;
    float2 uv : TEXCOORD0;
    float time : TEXCOORD1;
};

struct VertexOutput {
    float4 position : SV_Position;
    float2 uv : TEXCOORD0;
    float time : TEXCOORD1;
};

[goldy_vertex]
VertexOutput vs_main(VertexInput input) {
    VertexOutput output;
    output.position = float4(input.position, 0.0, 1.0);
    output.uv = input.uv;
    output.time = input.time;
    return output;
}

float hash(float p) {
    return frac(sin(p * 127.1) * 43758.5453);
}

[goldy_fragment]
float4 fs_main(VertexOutput input) : SV_Target {
    float2 uv = (input.uv - 0.5) * 2.0;
    float t = input.time;
    float3 color = float3(0.0, 0.0, 0.02);
    
    int num_stars = 200;
    float speed = 0.3;
    
    for (int i = 0; i < num_stars; i++) {
        float fi = float(i);
        
        // Random angle for each star
        float angle = hash(fi) * 6.28318;
        // Random max distance (how far from center it can go)
        float max_dist = 0.3 + hash(fi + 50.0) * 1.2;
        
        // All stars cycle at same speed, just different phases
        float phase = hash(fi + 100.0);
        float cycle = frac(t * speed + phase);
        
        // Distance from center increases as cycle goes 0->1
        float dist = max_dist * cycle;
        
        float star_x = cos(angle) * dist;
        float star_y = sin(angle) * dist;
        
        float pixel_dist = length(uv - float2(star_x, star_y));
        
        // Size increases with distance (closer = bigger)
        float size = 0.002 + cycle * 0.015;
        
        // Brightness increases with distance
        float brightness = cycle * smoothstep(size, 0.0, pixel_dist);
        
        // Fade in at spawn
        float fade = smoothstep(0.0, 0.1, cycle);
        
        color += float3(0.9, 0.95, 1.0) * brightness * fade;
    }
    
    return float4(clamp(color, float3(0.0, 0.0, 0.0), float3(1.0, 1.0, 1.0)), 1.0);
}

Shared Helpers

The windowed examples share a few modules. They are not registered as [[example]] targets; each example pulls them in with mod common; and friends.

examples/common.rs

Run limits (GOLDY_EXAMPLE_TIMEOUT / EXAMPLE_TIMEOUT), the trailing FPS window used by the GOLDY_PERF line, hidden-window creation so the first frame is never a blank flash, and render_pipeline / render_pipeline_for_surface, which rebuild a pipeline against the current colour-target format.

Present and readback stay in each example (SurfaceExchange::bind_render_target / bind_destination and (&mut submission >> &present).take()?, or (&mut submission >> &readback).take::<u8>() when GOLDY_EXAMPLE_CAPTURE is set). CaptureDump is only the packed-RGBA file that scripts/record_example_captures.sh stitches with ffmpeg. Optional: GOLDY_EXAMPLE_CAPTURE_FRAMES (default 75), GOLDY_EXAMPLE_CAPTURE_FPS (default 15), GOLDY_EXAMPLE_CAPTURE_WIDTH / HEIGHT (default 640×480).

//! Shared helpers for interactive examples (run limits, perf reporting, book captures).
//!
//! Set `GOLDY_EXAMPLE_CAPTURE` to a raw-RGBA output path to render headlessly
//! into pixels that `scripts/record_example_captures.sh` stitches with ffmpeg.
//! Optional: `GOLDY_EXAMPLE_CAPTURE_FRAMES` (default 75), `GOLDY_EXAMPLE_CAPTURE_FPS`
//! (default 15), `GOLDY_EXAMPLE_CAPTURE_WIDTH` / `HEIGHT` (default 640×480).
//!
//! Capture is file I/O plus a retained RGBA texture. Present still goes through
//! [`SurfaceExchange`] / [`MemoryExchange`] in each example.

use goldy::{
    RenderPipeline, RenderPipelineDesc, Runtime, ShaderModule, SurfaceExchange, Texture, TextureFlags, TextureFormat,
    TextureKind,
};
use std::fs::File;
use std::io::{BufWriter, Write};
use std::path::PathBuf;
use std::time::{Duration, Instant};
use winit::event_loop::ActiveEventLoop;
use winit::window::{Window, WindowAttributes};

/// Window attributes for examples that reveal only after the first frame is ready.
#[allow(dead_code)]
pub fn hidden_window(title: impl Into<String>, width: u32, height: u32) -> WindowAttributes {
    Window::default_attributes()
        .with_title(title.into())
        .with_inner_size(winit::dpi::LogicalSize::new(width, height))
        .with_visible(false)
}

/// Show a window after GPU init and an initial present path have completed.
pub fn reveal_window(window: &Window) {
    window.set_visible(true);
}

/// Rolling frame timestamps for windowed FPS (e.g. last 5s at exit).
#[allow(dead_code)]
pub struct FpsWindow {
    window: Duration,
    frames: Vec<Instant>,
}

#[allow(dead_code)]
impl FpsWindow {
    pub fn new(window_secs: f64) -> Self {
        Self {
            window: Duration::from_secs_f64(window_secs),
            frames: Vec::new(),
        }
    }

    pub fn record(&mut self, now: Instant) {
        self.prune(now);
        self.frames.push(now);
    }

    fn prune(&mut self, now: Instant) {
        let cutoff = now.checked_sub(self.window).unwrap_or(now);
        let keep_from = self.frames.partition_point(|t| *t < cutoff);
        if keep_from > 0 {
            self.frames.drain(..keep_from);
        }
    }

    /// Returns `(frames_in_window, window_span_secs, fps)` for the trailing window.
    pub fn stats(&mut self, now: Instant) -> Option<(u64, f64, f64)> {
        self.prune(now);
        let n = self.frames.len();
        if n == 0 {
            return None;
        }
        let span = now.duration_since(self.frames[0]).as_secs_f64();
        if span <= 0.0 {
            return None;
        }
        Some((n as u64, span, n as f64 / span))
    }
}

/// Build or rebuild a render pipeline for a colour-target format.
#[allow(dead_code)]
pub fn render_pipeline(
    device: &Runtime,
    shader: &ShaderModule,
    format: TextureFormat,
    desc: RenderPipelineDesc,
) -> anyhow::Result<RenderPipeline> {
    RenderPipeline::new(
        device,
        shader,
        shader,
        &RenderPipelineDesc {
            target_format: format,
            ..desc
        },
    )
}

/// Build or rebuild a render pipeline using the surface's current format.
#[allow(dead_code)]
pub fn render_pipeline_for_surface(
    device: &Runtime,
    shader: &ShaderModule,
    surface: &SurfaceExchange,
    desc: RenderPipelineDesc,
) -> anyhow::Result<RenderPipeline> {
    render_pipeline(device, shader, surface.format(), desc)
}

/// True when this process should dump frames instead of opening a window.
pub fn capture_requested() -> bool {
    match std::env::var("GOLDY_EXAMPLE_CAPTURE") {
        Ok(path) => !path.is_empty(),
        Err(_) => false,
    }
}

fn env_u32(key: &str, default: u32) -> u32 {
    std::env::var(key)
        .ok()
        .and_then(|raw| raw.parse().ok())
        .filter(|n| *n > 0)
        .unwrap_or(default)
}

fn env_f32(key: &str, default: f32) -> f32 {
    std::env::var(key)
        .ok()
        .and_then(|raw| raw.parse().ok())
        .filter(|n| *n > 0.0)
        .unwrap_or(default)
}

/// Packed-RGBA dump for mdBook clips (`GOLDY_EXAMPLE_CAPTURE`). Not a Goldy type.
pub struct CaptureDump {
    path: PathBuf,
    writer: Option<BufWriter<File>>,
    last: Option<Vec<u8>>,
    fps: f32,
    frames: u32,
    written: u32,
    width: u32,
    height: u32,
}

#[allow(dead_code)]
impl CaptureDump {
    fn capture_size() -> (u32, u32, u32, f32) {
        (
            env_u32("GOLDY_EXAMPLE_CAPTURE_WIDTH", 640),
            env_u32("GOLDY_EXAMPLE_CAPTURE_HEIGHT", 480),
            env_u32("GOLDY_EXAMPLE_CAPTURE_FRAMES", 75),
            env_f32("GOLDY_EXAMPLE_CAPTURE_FPS", 15.0),
        )
    }

    /// File dump from `GOLDY_EXAMPLE_CAPTURE`.
    pub fn from_env() -> anyhow::Result<Self> {
        let path = std::env::var("GOLDY_EXAMPLE_CAPTURE").expect("GOLDY_EXAMPLE_CAPTURE");
        let path = PathBuf::from(path);
        if let Some(parent) = path.parent() {
            if !parent.as_os_str().is_empty() {
                std::fs::create_dir_all(parent)?;
            }
        }
        let (width, height, frames, fps) = Self::capture_size();
        let file = File::create(&path)?;
        Ok(Self {
            path,
            writer: Some(BufWriter::new(file)),
            last: None,
            fps,
            frames,
            written: 0,
            width,
            height,
        })
    }

    /// In-memory frames only (e.g. `multi_window` panels before hstack).
    pub fn memory(width: u32, height: u32) -> Self {
        let fps = env_f32("GOLDY_EXAMPLE_CAPTURE_FPS", 15.0);
        Self {
            path: PathBuf::new(),
            writer: None,
            last: None,
            fps,
            frames: u32::MAX,
            written: 0,
            width,
            height,
        }
    }

    pub fn width(&self) -> u32 {
        self.width
    }

    pub fn height(&self) -> u32 {
        self.height
    }

    pub fn size(&self) -> (u32, u32) {
        (self.width, self.height)
    }

    pub fn format() -> TextureFormat {
        TextureFormat::Rgba8Unorm
    }

    /// Virtual clock `written / fps` so clips are deterministic.
    pub fn time(&self) -> f32 {
        self.written as f32 / self.fps
    }

    pub fn dt(&self) -> f32 {
        1.0 / self.fps
    }

    pub fn finished(&self) -> bool {
        self.written >= self.frames
    }

    pub fn write_rgba(&mut self, pixels: &[u8]) -> anyhow::Result<()> {
        let expected = (self.width as usize)
            .saturating_mul(self.height as usize)
            .saturating_mul(4);
        anyhow::ensure!(
            pixels.len() == expected,
            "capture frame is {} bytes, expected {expected} ({}x{} rgba)",
            pixels.len(),
            self.width,
            self.height
        );
        if let Some(file) = &mut self.writer {
            file.write_all(pixels)?;
        }
        self.last = Some(pixels.to_vec());
        self.written += 1;
        if let Some(file) = &mut self.writer {
            if self.written >= self.frames {
                file.flush()?;
                println!(
                    "GOLDY_CAPTURE: wrote {} frames ({}x{} rgba) to {}",
                    self.written,
                    self.width,
                    self.height,
                    self.path.display()
                );
            }
        }
        Ok(())
    }

    pub fn take_rgba(&mut self) -> Option<Vec<u8>> {
        self.last.take()
    }
}

/// Retained RGBA8 texture for `copy_to_texture` plus a host claim on the capture path.
#[allow(dead_code)]
pub fn capture_readback(device: &Runtime, width: u32, height: u32) -> anyhow::Result<Texture> {
    device.acquire_texture(
        width,
        height,
        TextureFormat::Rgba8Unorm,
        TextureKind::Direct,
        TextureFlags::COPY_SRC | TextureFlags::COPY_DST,
        None,
    )
}

/// Horizontal concat of equal-sized packed RGBA panels (for `multi_window` capture).
#[allow(dead_code)]
pub fn hstack_rgba(panels: &[&[u8]], width: u32, height: u32) -> anyhow::Result<Vec<u8>> {
    let row_bytes = width as usize * 4;
    let panel_bytes = row_bytes * height as usize;
    for (i, panel) in panels.iter().enumerate() {
        anyhow::ensure!(
            panel.len() == panel_bytes,
            "hstack panel {i} is {} bytes, expected {panel_bytes}",
            panel.len()
        );
    }
    let n = panels.len();
    let mut out = vec![0u8; panel_bytes.saturating_mul(n)];
    for y in 0..height as usize {
        for (i, panel) in panels.iter().enumerate() {
            let src = y * row_bytes;
            let dst = y * row_bytes * n + i * row_bytes;
            out[dst..dst + row_bytes].copy_from_slice(&panel[src..src + row_bytes]);
        }
    }
    Ok(out)
}

/// Run limit in seconds from `GOLDY_EXAMPLE_TIMEOUT` or `EXAMPLE_TIMEOUT`.
pub fn run_limit_secs() -> Option<f64> {
    for key in ["GOLDY_EXAMPLE_TIMEOUT", "EXAMPLE_TIMEOUT"] {
        if let Ok(raw) = std::env::var(key) {
            if let Ok(secs) = raw.parse::<f64>() {
                if secs > 0.0 {
                    return Some(secs);
                }
            }
        }
    }
    None
}

/// Exit the event loop once the run limit elapses so `Drop` can print `GOLDY_PERF`.
pub fn exit_if_timed_out(event_loop: &ActiveEventLoop, start: Instant) {
    if let Some(limit) = run_limit_secs() {
        if start.elapsed() >= Duration::from_secs_f64(limit) {
            event_loop.exit();
        }
    }
}

examples/digital_clock_shared.rs

Seven-segment digit geometry for digital_clock.

//! Shared digital-clock rendering helpers for the `digital_clock` example.

use goldy::types::Color;

/// Vertex with 2D position and RGBA color.
#[goldy::gpu]
#[derive(Debug)]
pub struct ClockVertex {
    pub position: [f32; 2],
    pub color: [f32; 4],
}

impl ClockVertex {
    pub const fn new(x: f32, y: f32, color: Color) -> Self {
        Self {
            position: [x, y],
            color: [color.r, color.g, color.b, color.a],
        }
    }
}

/// Seven-segment display patterns.
/// Order: top, top-left, top-right, middle, bottom-left, bottom-right, bottom
pub const SEGMENT_PATTERNS: [[bool; 7]; 11] = [
    [true, true, true, false, true, true, true],       // 0
    [false, false, true, false, false, true, false],   // 1
    [true, false, true, true, true, false, true],      // 2
    [true, false, true, true, false, true, true],      // 3
    [false, true, true, true, false, true, false],     // 4
    [true, true, false, true, false, true, true],      // 5
    [true, true, false, true, true, true, true],       // 6
    [true, false, true, false, false, true, false],    // 7
    [true, true, true, true, true, true, true],        // 8
    [true, true, true, true, false, true, true],       // 9
    [false, false, false, false, false, false, false], // 10 = blank (for colon position)
];

pub const COLORS: [Color; 8] = [
    Color {
        r: 0.2,
        g: 1.0,
        b: 0.3,
        a: 1.0,
    },
    Color {
        r: 1.0,
        g: 0.3,
        b: 0.2,
        a: 1.0,
    },
    Color {
        r: 1.0,
        g: 0.6,
        b: 0.0,
        a: 1.0,
    },
    Color {
        r: 1.0,
        g: 1.0,
        b: 0.2,
        a: 1.0,
    },
    Color {
        r: 0.2,
        g: 1.0,
        b: 1.0,
        a: 1.0,
    },
    Color {
        r: 0.4,
        g: 0.6,
        b: 1.0,
        a: 1.0,
    },
    Color {
        r: 0.8,
        g: 0.3,
        b: 1.0,
        a: 1.0,
    },
    Color {
        r: 1.0,
        g: 0.4,
        b: 0.8,
        a: 1.0,
    },
];

pub fn quad_vertices(x: f32, y: f32, w: f32, h: f32, color: Color) -> [ClockVertex; 6] {
    [
        ClockVertex::new(x, y, color),
        ClockVertex::new(x + w, y, color),
        ClockVertex::new(x + w, y + h, color),
        ClockVertex::new(x, y, color),
        ClockVertex::new(x + w, y + h, color),
        ClockVertex::new(x, y + h, color),
    ]
}

pub fn pixel_to_ndc(px: f32, py: f32, width: f32, height: f32) -> (f32, f32) {
    let x = (px / width) * 2.0 - 1.0;
    let y = 1.0 - (py / height) * 2.0;
    (x, y)
}

pub fn digit_vertices(
    digit: u8,
    cx: f32,
    cy: f32,
    scale: f32,
    color: Color,
    width: f32,
    height: f32,
) -> Vec<ClockVertex> {
    let mut vertices = Vec::new();

    let seg_w = 60.0 * scale;
    let seg_h = 12.0 * scale;
    let dig_h = 120.0 * scale;
    let gap = 4.0 * scale;

    if digit == 10 {
        let dot_size = seg_h * 1.5;
        let dot_spacing = dig_h * 0.5;

        let (x, y) = pixel_to_ndc(cx - dot_size / 2.0, cy - dot_spacing - dot_size / 2.0, width, height);
        let (w, h) = (dot_size / width * 2.0, dot_size / height * 2.0);
        vertices.extend_from_slice(&quad_vertices(x, y, w, -h, color));

        let (x, y) = pixel_to_ndc(cx - dot_size / 2.0, cy + dot_spacing - dot_size / 2.0, width, height);
        vertices.extend_from_slice(&quad_vertices(x, y, w, -h, color));

        return vertices;
    }

    let pattern = SEGMENT_PATTERNS[digit as usize];

    let mut add_segment = |px: f32, py: f32, pw: f32, ph: f32| {
        let (x, y) = pixel_to_ndc(px, py, width, height);
        let (w, h) = (pw / width * 2.0, ph / height * 2.0);
        vertices.extend_from_slice(&quad_vertices(x, y, w, -h, color));
    };

    if pattern[0] {
        add_segment(cx - seg_w / 2.0, cy - dig_h, seg_w, seg_h);
    }
    if pattern[1] {
        add_segment(
            cx - seg_w / 2.0 - seg_h,
            cy - dig_h + seg_h + gap,
            seg_h,
            dig_h - seg_h - gap * 2.0,
        );
    }
    if pattern[2] {
        add_segment(
            cx + seg_w / 2.0,
            cy - dig_h + seg_h + gap,
            seg_h,
            dig_h - seg_h - gap * 2.0,
        );
    }
    if pattern[3] {
        add_segment(cx - seg_w / 2.0, cy - seg_h / 2.0, seg_w, seg_h);
    }
    if pattern[4] {
        add_segment(cx - seg_w / 2.0 - seg_h, cy + gap, seg_h, dig_h - seg_h - gap * 2.0);
    }
    if pattern[5] {
        add_segment(cx + seg_w / 2.0, cy + gap, seg_h, dig_h - seg_h - gap * 2.0);
    }
    if pattern[6] {
        add_segment(cx - seg_w / 2.0, cy + dig_h - seg_h, seg_w, seg_h);
    }

    vertices
}

#[derive(Debug, Clone, Copy, Default)]
pub struct TimeData {
    pub hours: u8,
    pub minutes: u8,
    pub seconds: u8,
}

impl TimeData {
    pub fn from_elapsed_secs(elapsed: u64) -> Self {
        Self {
            hours: ((elapsed / 3600) % 100) as u8,
            minutes: ((elapsed % 3600) / 60) as u8,
            seconds: (elapsed % 60) as u8,
        }
    }

    pub fn to_digits(&self) -> [u8; 8] {
        [
            self.hours / 10,
            self.hours % 10,
            10,
            self.minutes / 10,
            self.minutes % 10,
            10,
            self.seconds / 10,
            self.seconds % 10,
        ]
    }
}

pub fn generate_clock_vertices(time: TimeData, color: Color, width: u32, height: u32) -> Vec<ClockVertex> {
    let digits = time.to_digits();

    let scale = height as f32 / 720.0;
    let digit_width = 80.0 * scale;
    let colon_width = 40.0 * scale;
    let spacing = 20.0 * scale;

    let total_width = digit_width * 6.0 + colon_width * 2.0 + spacing * 7.0;

    let cy = height as f32 / 2.0;
    let mut cx = (width as f32 - total_width) / 2.0 + digit_width / 2.0;

    let mut all_vertices = Vec::new();

    for &digit in digits.iter() {
        let w = if digit == 10 { colon_width } else { digit_width };
        let verts = digit_vertices(digit, cx, cy, scale, color, width as f32, height as f32);
        all_vertices.extend_from_slice(&verts);
        cx += w + spacing;
    }

    all_vertices
}

#[allow(dead_code)]
#[derive(Debug, Clone, Default)]
pub struct ClockState {
    pub color_index: usize,
    pub paused: bool,
    pub accumulated_secs: u64,
}

#[allow(dead_code)]
impl ClockState {
    pub fn color(&self) -> Color {
        let mut color = COLORS[self.color_index];
        if self.paused {
            color.r *= 0.5;
            color.g *= 0.5;
            color.b *= 0.5;
        }
        color
    }

    pub fn background_color(&self) -> Color {
        let bg = if self.paused { 0.06 } else { 0.02 };
        Color {
            r: bg,
            g: bg,
            b: bg,
            a: 1.0,
        }
    }

    pub fn next_color(&mut self) {
        self.color_index = (self.color_index + 1) % COLORS.len();
    }

    pub fn toggle_pause(&mut self, current_elapsed: u64) {
        if self.paused {
            self.paused = false;
        } else {
            self.accumulated_secs = current_elapsed;
            self.paused = true;
        }
    }
}

examples/instance2d.rs

The per-instance struct for instancing, laid out to match QuadInstance in instancing_update.slang and instancing_render.slang.

//! Per-instance data for the instancing example.

#[goldy::gpu]
#[derive(Debug, Default)]
pub struct Instance2D {
    pub position: [f32; 2],
    pub rotation: f32,
    pub scale: f32,
    pub color: [f32; 4],
}

impl Instance2D {
    pub const fn new(x: f32, y: f32, rotation: f32, scale: f32, color: [f32; 4]) -> Self {
        Self {
            position: [x, y],
            rotation,
            scale,
            color,
        }
    }
}

Motivation

Goldy implements the Fondaco Machine on modern GPUs: programs describe parcels and schemes; the runtime manages the physical medium, derives precedences from ownership, mediates present through exchanges, and realizes host reads as public claims. For the normative machine spec see the Machine Specification; for what Goldy ships today see the runtime mapping.

The Problem with "Modern" Graphics APIs

DX12, Vulkan, and Metal are commonly called modern APIs, but they were designed over a decade ago for hardware that has since changed dramatically. The GPU architectures those APIs targeted lacked coherent caches, bindless descriptors, and 64-bit pointers. The APIs compensated with layers of indirection — descriptor sets, render pass objects, explicit image layout transitions, pipeline layouts as first-class objects — that served as hints and contracts for hardware that needed them.

Furthermore, high-performance GPU programs tend to converge on the same shape, whether they are written with PyTorch, CUDA, Metal, Vulkan, or something else. They are not best understood as a stream of API calls. They are graphs whose nodes are kernels, copies, and foreign operations, and whose edges describe data dependencies. Independent nodes may run concurrently; dependent nodes require an ordering mechanism such as a barrier, event, semaphore, or stream dependency.

PyTorch makes this especially visible. A model is a graph of tensor operations; autograd constructs another graph for the backward pass, and compilers such as TorchInductor capture, specialize, fuse, and schedule that work. Beneath it, CUDA libraries submit kernels and transfers to streams, use events to express cross-stream dependencies, reuse temporary allocations according to tensor lifetimes, and fuse adjacent operations to reduce launch and memory-traffic costs. Hand-written CUDA programs eventually acquire the same machinery: stream graphs, memory pools, dependency tracking, and explicit synchronization around shared buffers.

Graphics workloads arrive at the same structure from another direction. A frame graph records render, compute, copy, and presentation passes together with how each pass reads or writes resources. From those declarations, an engine derives execution order, barriers, transient-memory aliasing, queue placement, and opportunities for overlap. The APIs differ, but the optimization problem is the same: preserve data dependencies while minimizing synchronization, allocation, launch, and memory-traffic costs.

This convergence suggests that the graph and its resource relationships are the durable program model. Descriptor updates, barriers, command buffers, streams, and semaphores - and yes, sometimes even CPU waits - are mechanisms a runtime can derive from that model for a particular GPU.

Yet every application using graphics APIs still pays the complexity cost of the old model, and do not idiomatically map to the model of the best GPU programs.

Why Bindless Matters

Traditional GPU programming organizes resources into descriptor sets — fixed layouts of bindings that must be declared ahead of time, allocated from pools, and swapped between draw calls. This model creates a cascade of complexity:

  • Pipeline layout explosion: Every unique combination of descriptor set layouts produces a distinct pipeline layout, and each pipeline layout dimension multiplies the total pipeline state permutation count.
  • CPU overhead: Updating and binding descriptor sets each frame is a significant portion of CPU-side draw call cost.
  • Shader inflexibility: Shaders are coupled to their binding layout; changing which resources a shader accesses means changing the pipeline.

Bindless resource access replaces all of this with a single concept: resources live in GPU-visible memory, and shaders access them by index. There are no set layouts to declare, no pools to manage, no binding points to track. A shader that needs buffer #7 just reads slot 7 from a flat descriptor heap.

This isn't exotic — it's how game engines have been working internally for years. Goldy makes it the public API rather than hiding it behind compatibility abstractions.

Why a Dependency Graph (Scheme)

Bindless access means shaders can read any resource at any time. The traditional model of inserting barriers at the call site ("I'm about to read this buffer, so transition it now") breaks down when the set of resources a dispatch touches isn't known until the shader runs.

Goldy uses a retained scheme — a dependency graph you record once and submit many times — to solve this. You declare nodes and their resource dependencies; Goldy derives the barriers, layout transitions, and execution order automatically. This is both safer (no missed barriers) and simpler (no manual synchronization) than the alternative.

The scheme also enables Goldy to batch and reorder work across the frame, which matters for compute-heavy workloads where multiple dispatches feed into each other before anything reaches the screen.

Why Slang

The shader language landscape is fragmented. GLSL, HLSL, MSL, and WGSL each target a subset of platforms, and none is a clean superset of the others. Libraries that support multiple shading languages maintain translation layers and per-language workarounds, which is a significant source of bugs and complexity.

Slang solves this at the source level. A single Slang source file compiles to SPIR-V (Vulkan), DXIL (DX12), and MSL (Metal). It uses HLSL-familiar syntax with additions that matter for modern GPU programming:

FeatureWhy it matters
Modules and importTrue separate compilation, no #include fragility
GenericsType-safe reusable shader code
Automatic differentiationFirst-class for ML and physics workloads
Khronos governanceLong-term stability and active development

By committing to Slang as the sole shader language, Goldy eliminates an entire category of cross-platform bugs and keeps its codebase focused on GPU work rather than shader translation. By embedding a verified compatible version of slang, goldy simplifies packaging.

What Goldy Sheds

Goldy's bindless model and modern-hardware baseline make several traditional GPU programming concepts unnecessary. These aren't missing features — they're intentional design choices that keep the API small and the programming model coherent.

No Descriptor Set Management

Traditional APIs require you to declare descriptor set layouts, allocate descriptor pools, write descriptor sets, and bind them before each draw or dispatch. A typical Vulkan pipeline touches three to four descriptor set objects before anything reaches the GPU.

Goldy replaces all of this with a flat bindless heap. Resources get a slot index when created, and shaders access them by that index. There are no layouts, no pools, no binding calls.

// Shader receives resources by index — no descriptor sets
[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Particle> particles, ThreadId id) {
    particles[id.x].position += particles[id.x].velocity;
}

This also eliminates pipeline layouts as objects. In Vulkan, each unique combination of descriptor set layouts produces a pipeline layout, which is baked into the pipeline at creation time. Goldy's single global bindless layout means one pipeline layout for all pipelines.

No Manual Barrier Insertion

In Vulkan and DX12, you manually insert memory barriers and image layout transitions to tell the GPU when a resource changes from "written by compute" to "read by fragment" (or any other transition). Missing a barrier is a silent correctness bug; inserting too many is a performance bug.

Goldy's scheme (dependency graph) handles this automatically. You declare what each node reads and writes; Goldy derives the minimal set of barriers and transitions. This is both safer and typically more efficient than hand-placed barriers, because the scheme has a global view of the frame.

No Shader Permutation Systems

Traditional engines maintain thousands of shader variants — combinations of feature flags, render pass compatibility, descriptor set layout versions, and pipeline state. Some ship dedicated cloud infrastructure just to compile and cache them all.

Goldy collapses most of the dimensions that drive permutation counts:

Traditional dimensionGoldy equivalent
Render pass compatibilityDynamic rendering — no render pass objects
Descriptor set layoutOne global bindless layout
Pipeline layoutImplicit from the global layout
Viewport/scissor stateDynamic state, not baked into PSO

What remains — shader source × vertex format × target format × depth config — is a small, manageable space. Goldy addresses pipeline variety by having fewer pipelines, not by building infrastructure to manage many variants.

That stance is compatible with compiling a predicted specialized compute program. When a dispatch's with_param scalars hold still, the runtime bakes those words into a second pipeline for that site — at most one specialized program, plus the universal program that is always correct. It does not enumerate a permutation matrix, and it does not ask the caller to name flags. The worthwhile cases are the ones that gate whole code paths behind a scalar (a tint, an AA mode, a filter), not the ones that only feed arithmetic; see What baking actually compiles.

Minimal Pipeline State Management

A Vulkan VkGraphicsPipelineCreateInfo touches blend state, depth/stencil state, rasterizer state, multisample state, input assembly, viewport/scissor, dynamic state flags, render pass, subpass, pipeline layout, and shader stages. Many of these are baked in at pipeline creation time, producing the combinatorial explosion that drives PSO caches.

Goldy uses dynamic rendering and dynamic state to move viewport, scissor, and render target configuration out of the pipeline object. The remaining pipeline state is intentionally minimal:

#![allow(unused)]
fn main() {
let pipeline = RenderPipeline::new(&device, &shader, &shader, &desc)?;
}

Blend mode, depth testing, and vertex format are still part of the pipeline — they represent genuine hardware configuration. But the many compatibility dimensions that traditional APIs bake in are gone.

No Separate Compute API

OpenCL introduced compute to GPUs as an entirely separate API with its own device model, memory model, and dispatch semantics. Even "unified" APIs like Vulkan treat compute as a second-class citizen — compute pipelines and graphics pipelines share almost no code paths.

In Goldy, compute is a first-class citizen on the same footing as graphics. Compute shaders use the same bindless resource model, the same buffer types, and the same scheme. A compute dispatch that writes to a buffer and a render pass node that reads from it are peers in the graph.

#![allow(unused)]
fn main() {
// Compute updates particles, render draws them — same scheme
scheme.node("update", &compute_pipeline)
    .with_parcel(&particle_buf, NodeAccess::ReadWrite)
    .dispatch(workgroups, 1, 1);
let mut pass = scheme.render_pass("draw", &scene_rt);
pass.with_parcel(&particle_buf, NodeAccess::Read);
// ...
}

The Design Principle

Each of these omissions follows the same logic: if modern hardware doesn't need a concept for correctness or performance, Goldy doesn't expose it. The result is an API where the concepts that remain — buffers, textures, shaders, pipelines, scheme — each carry their weight.

Shader Specialization Prediction

Programs know things about a frame that a shader cannot see: whether any image in the scene carries a tint, which antialiasing path the surface needs, whether a filter is active, what the output format is. Each of those facts could pick a smaller, faster GPU program — but only if something decides, before a retained scheme is recorded, which program the next hundred frames will want.

This note describes how Goldy makes that decision generically, without growing a named feature flag per specialization and without caching the full combinatorial set of variants.

Specialization is an implementation detail of the runtime, not an API. Every program that submits a retained scheme is a potential beneficiary without changing a line, and the only knob is an environment variable that turns the whole thing off. Baking is a full recompile: the specialized program can elide whole code paths, not merely load a constant. How optimistic to be depends on whether the baked word gates work; see What baking actually compiles.

Status: implemented in src/specialization.rs, driven from Scheme::submit and Scheme::set_node_param. Sections below describe the shipped behaviour; the closing Follow-ups list what was deliberately left out.

Why not a flag per feature

The obvious approach is to let the caller name the fact — a has_tint boolean on a worker description, a second retained scheme for the tinted case — and it does not scale. Tint is one of many statically determinable specializations. Antialiasing mode, filters, non-default blend, and output format are others, and a higher-level compiler will emit facts Goldy has never heard of. Each named flag doubles the number of retained schemes a client has to hold, and each retained scheme carries its own command lists, topology edges, and partition bookkeeping.

Goldy already resists shader permutations rather than industrializing them (see What Goldy Sheds). The mechanism here follows that stance: the runtime holds at most one specialized program per dispatch site, plus a universal program that is always correct.

A CPU branch predictor is the precedent. It does not expose a named flag for each if in user code, and it does not ask the program what it expects; it indexes opaque history by instruction identity and outcome.

CPU branch predictorGoldy
Branch PCStable dispatch-site identity (NodeId on a scheme)
Branch outcomeScalar param wire words the site last dispatched with
Hidden history tableSmall per-site predictor state
Generic code pathUniversal shader (reads its params at runtime)
Optimized targetSpecialized PSO plus the re-recorded partition that binds it

Where the facts come from

Goldy does not need to be told what to specialize on, because it already holds the facts: the scalar params of every dispatch site.

Scheme::with_param takes a u32 wire word. Those words are the program's own encoding of whatever it decided — a tint factor, a mode enum, a filter toggle, a count — and Goldy stores them on the dispatch node, hashes them into the emission fingerprint, and writes them into push constants on every record. A value that has been the same for the last ten frames is a fact about the scene, and the runtime can see that without being told.

So the specialization key for a site is derived, not supplied: it is the tuple of scalar wire words the site dispatches with. Goldy never interprets a word. It only compares words for equality across frames, which is exactly what makes the mechanism generic — a new specialization axis is a new with_param in the caller, not a new field in a Goldy struct.

Baking a param

The virtual-main transform reads every scalar param through a preprocessor macro that defaults to the push-constant word:

#ifndef _GOLDY_SPEC_CS_MAIN_UW0
#define _GOLDY_SPEC_CS_MAIN_UW0 _uw0
#endif

[shader("compute")]
[numthreads(64, 1, 1)]
void cs_main(uniform uint _bw0, /* ... */ uniform uint _uw0, /* ... */) {
    uint factor = _GOLDY_SPEC_CS_MAIN_UW0;   // was: _uw0
    // ...
}

Defining that macro to a wire-word literal bakes the value. ShaderModule::variant already merges defines onto a retained module, so a specialized program is variant(&[("_GOLDY_SPEC_CS_MAIN_UW0", "10u")]) and nothing more. The macro is named after the author's [goldy_compute] function (cs_main above; _GOLDY_SPEC_TINT_UW0 for void tint(...)) so that sources with several entries specialize independently, and the define always carries the raw u32 wire word, so the runtime needs no knowledge of the param's Slang type — the existing asfloat / asint / != 0u decode applies to the literal exactly as it applied to the push-constant word. The runtime recovers the function name from the retained source (single_compute_entry_name); a source with zero or several compute entries has no unambiguous macro to define and is left alone.

Two properties make this safe to swap under an already-recorded dispatch:

  • The parameter ABI does not move. The wrapper's signature always declares all eight _uw words regardless of use, so baking one leaves the others at their original offsets. A site with a baked factor and a runtime bias still reads bias correctly. The WebGPU and CUDA lowerings keep their uniform field and kernel argument for the same reason.
  • The recorded payload does not change. Baking rewrites the program, not the command list. The same wire words stay baked into the partition; the specialized program simply stops reading them.

What baking must not do is change the pipeline's binding layout, which is a real constraint on some backends and is covered below.

What baking actually compiles

Baking is a full recompile, not a constant patch. The specialized program is a new ShaderModule::variant with the _GOLDY_SPEC_* macros defined to wire-word literals, so Slang, the backend IR (SPIR-V / DXIL / MSL / PTX), and the driver's ISA compiler all see a compile-time constant. That constant can participate in folding anywhere it flows: branch conditions, loop trip counts, array indices, asfloat(...) arithmetic, resource selection. Nothing is rewritten at bind time, and the universal program is left alone.

Elision is split across two compilers, and that split is easy to misread.

Goldy's default OptimizationLevel::Default does not run Slang's dead-code pass, so a branch whose condition has become false still appears in the dumped SPIR-V — as OpConstantFalse plus the unreachable block. The push-constant load is already gone at that point. Folding a constant condition, deleting the dead block, and reallocating registers is the first thing any production driver compiler does (NVIDIA, AMD's LLVM, Apple's, DXC, Mesa lavapipe). Dumping SPIR-V instruction counts therefore understates the specialized program: the variant that looks almost the same as the universal one at the SPIR-V level can be a fraction of the size after the driver.

Raising Slang to OptimizationLevel::Maximal does the DCE in SPIR-V itself. That moves compile time; it does not improve the ISA the GPU runs. Variants inherit the module's optimization level and there is no reason to raise it for baking.

Measured on lavapipe through Goldy's real Vulkan compile path, a compute shader whose has_tint scalar gated a tinting loop:

ProgramSPIR-V instructionsSPIR-V still has the branch+loopDriver NIR instructionsif / fmul
Universal212yes2011 / 55
Baked has_tint = 0210yes, on OpConstantFalse360 / 0
Baked has_tint = 1207yes1940 / 55
Baked has_tint = 0, Slang Maximal91no360 / 0

The specialized-off program dropped the extra buffer loads, the extra push-constant loads, the unrolled loop, and the registers they held. The specialized-on program dropped only the if and the two unused push-constant loads, which is the expected "only what the constant makes dead goes away" behaviour. Numbers will differ on a hardware compiler; the split between Slang and the driver will not.

When this is worth expecting

A uniform branch on a push constant is already non-divergent and cheap. Removing the if itself is not the win. The win is whatever sits behind it.

Kind of with_paramWhat baking doesHow optimistic to be
Feature / mode flag that gates whole code paths (has_tint, AA mode, filter on)The driver deletes the unused path, its memory traffic, and the registers it held — occupancy can riseHigh, if the gated work is real
Loop bound or "which of N resources" indexTrip count or index becomes a literal; unrolling and addressing followHigh when the bound is small and the body is heavy
Arithmetic coefficient (x * factor + bias)One less push-constant load; maybe a strength reductionModest
Value that changes every few framesPredictor raises the bake threshold and leaves the slot dynamicNone — it will not stay promoted
Value that lives in a bound bufferInvisible to the predictorNone — put scene facts in with_param if they should bake

The cost is a full Slang compile plus PSO creation on a worker thread, which is why warm waits for two clean hits and promotion waits for ten. A scheme whose many nodes all stabilize at once will spend real CPU in the background for a moment; the 16-entry per-scheme cache bounds the steady state.

Shader authors do not opt sites in, but they do choose where facts live. A mode flag in a Scattered buffer cannot bake. A mode flag passed with with_param, with the expensive work behind if (has_tint != 0u), is exactly the shape the predictor was built for. That is also why Goldy still resists permutation systems: the runtime holds at most one specialized program per dispatch site, not the cross product of every flag.

Per-slot, not per-tuple

Stability is tracked per param slot, not over the whole tuple. A site that passes a stable mode flag and a frame counter can still bake the flag. Keying on the whole tuple would give up on any site with a single volatile param, which in practice is most of them.

Why with_param and not a bound buffer

A fact in a buffer is invisible: Goldy sees a parcel read, not a value, and it would have to read back GPU memory to learn anything. A fact in a with_param is already on the CPU, already hashed, and already known to be stable or not.

The apparent objection is that scalar params are part of the partition retention fingerprint, so flipping one re-records. That is not a cost of this mechanism, it is the signal that drives it: a param that changes every frame re-records anyway and will never earn promotion, and a param that is stable is both free to retain and profitable to bake.

Two caches, deliberately separate

CacheKeyed byHoldsEvicting it costs
Variant PSO cache(shader identity, baked slots and values)Compiled pipelineA recompile
Per-site predictionDispatch-site identityWhich slots are bakedA re-record

Keeping them separate means a site can be demoted without throwing away the compiled pipeline, so re-promotion later is nearly free: a demoted site whose words come back is promoted straight from the cache once its streak recovers, with no compile.

The PSO cache is scheme-scoped and bounded (16 variants, LRU). A device-scoped cache was the first design and does not work as stated: a ComputePipeline holds a strong Runtime, so a cache inside the device would form a reference cycle and the device would never be dropped. Scoping the cache to the scheme also settles ownership — the scheme holds an Arc to the variant it has promoted, and the cache holds another, so eviction from the cache never destroys a pipeline a node still binds. The price is that two schemes running the same shader with the same stable words each compile their own variant; the Slang disk cache and the driver's PSO cache make the second compile cheap, and device-level sharing is listed as a follow-up.

Compile workers write into the cache directly, so a variant that finishes after its site lost interest still lands there rather than being dropped.

The Slang bytecode disk cache already sits underneath both, keyed on post-transform source plus defines, so a cold process still avoids full recompiles.

The predictor

The hard part is not compiling two pipelines. It is first-frame uncertainty: when a scheme is recorded, nothing knows whether the next hundred frames will use the same params. So the predictor never speculates about the current frame — the params for the current frame are already known exactly, and history only decides whether to prepare a specialization for future frames.

Per-slot streaks

Every dispatch node with at least one with_param gets a site record when it is declared (Scheme::node, ComputeNodeRecord::commit_dispatch_scheme, or a caller-side set_node_pipeline, which starts the site over against the new pipeline). The record holds the caller's pipeline (the universal), the shader's provenance, the words seen at the last submit, and per slot:

  • a streak — consecutive clean submits during which the word held its value. A changed word resets its slot to zero on any submit; an unchanged word advances only when the scheme was otherwise clean, so a scheme that re-records every frame for unrelated reasons keeps its history but does not earn promotions from it. Topology dirtiness (a foreign scheme changing shared-parcel interaction) resets every slot.
  • a bake threshold — the streak the slot needs before it is baked. It starts at the warm threshold and grows every time the slot invalidates a compile or a promotion (to the promote threshold, then doubling), so a fact that flips every few frames is baked once, disproved once, and thereafter left dynamic while its neighbours specialize.

The set of slots at or past their threshold is the site's bake target. It is per-slot: a site with a stable mode flag and a moving counter bakes the flag.

Stages

Compilation is expensive and uncancellable once started, so promotion is staged. Compiling is separated from swapping, with different thresholds:

StageEntered whenBehaviour
ObservingSite declared, or after a demotionUniversal runs; streaks accumulate
WarmingBake target non-empty and not what is already promotedUniversal still runs; a variant baking the target compiles on a worker thread (or is taken from the cache)
ReadyThe compile landedUniversal still runs until every baked slot's streak reaches the promote threshold
PromotedEvery baked slot at or past the promote thresholdNode rebound to the variant as a params-only re-record; the site keeps observing the slots it did not bake and may warm a wider variant
PinnedThree failed compilesUniversal only; the site is never consulted again

Defaults are warm at 2, promote at 10 (SpecializationPolicy). The two-hit warm threshold skips the common ping-pong case, where a value alternates every frame and no specialization would ever pay off. The ten-hit promotion threshold buys confidence that the streak is a scene property rather than a coincidence, and it usually gives the compile enough time to finish before the swap is wanted.

The predictor runs at the top of Scheme::submit, before dirtiness is read for recording, so a promotion is recorded by the very submit that decided it and the scheme is clean again afterwards. It never touches a node whose current words differ from the variant's baked words — a swap is only ever to a program that agrees with the frame.

Demotion is mandatory, not optional

A promoted site is only correct while its baked words match what the program is dispatching. Scheme::set_node_param is the single place a param can change, so it is also the place that must un-promote: if the new word differs from the baked one, the site's pipeline reverts to universal in the same call that marks the scheme params-dirty. The partition then re-records with the universal program bound, and no submit ever runs a program whose baked value disagrees with the frame.

This is the one invariant the implementation cannot get wrong, and it is why the swap lives inside the runtime rather than in a caller's hands.

Cancellation is best-effort

A baked word changing while a compile is in flight — in set_node_param, or observed at the next submit — drops the job and raises its cancel flag. The flag is honoured before the compile starts; it cannot interrupt work already running, because Slang compilation runs behind a process-global lock and the driver's pipeline creation is not interruptible, so a cancelled compile may still run to completion.

When it does, the resulting pipeline is still inserted into the variant PSO cache. It is not swapped in, but it is not wasted either: if those words come back and earn promotion later, there is nothing left to compile.

Cost model

A streak alone is not the whole story. Recording a partition costs CPU time and, on backends that require it, a wait for the previous retained command storage to retire; a tiny dispatch cannot repay that. The shipped predictor has no dispatch-size term — the thresholds are the only cost model — so a small, long-lived, stable dispatch pays one re-record it may never recoup. That is bounded (one record per promotion, promotions are rare by construction) and is listed as a follow-up rather than guessed at without profiles.

First frame, oscillation, reset

  • First frames always run universal. There is no speculation before any history exists.
  • A miss demotes immediately, as above.
  • Oscillation is absorbed by the two-hit warm gate and the growing per-slot bake threshold, and three failed compiles move the site to Pinned so a pathological shader cannot make Goldy compile forever.
  • Reset happens on topology dirtiness (a foreign scheme changing shared-parcel interaction topology, which already forces a re-record) and on a caller-side set_node_pipeline, which replaces the universal. Both clear streaks; neither clears the PSO cache.
  • Turning the feature off at runtime (the environment variable is read every submit) demotes every promoted site on the next submit and drops in-flight compiles.

Baking must not change the binding layout

A specialized variant is swapped under bindings that were already resolved when the node was recorded, so it has to agree with the universal program about what those bindings are. Baking can break that agreement, because making a value constant can make a resource unreachable, and a compiler is entitled to stop expecting a binding nothing reads.

This is not hypothetical, and it is not uniform across backends. Vulkan, DX12, and Metal bind through a bindless heap, so a pipeline's expectations follow the shader signature and survive baking. WebGPU does not: wgpu derives the bind group layout from the compiled WGSL, so it reflects usage. Two consequences, both observed:

  • Baking a value that was the only reason a bound buffer was read drops that buffer from the layout, and the recorded dispatch then supplies a binding the pipeline no longer declares.
  • Baking every scalar of an entry point leaves the generated user-params uniform buffer unreferenced, which drops it the same way — so on those backends at least one scalar slot has to stay dynamic.

Neither constraint is enforced by the predictor, because the predictor does not run where they apply (next section). On the backends where it does run, the layout follows the signature and a variant binds exactly like its universal, so the site's bake target can be every stable slot.

A caller-side set_node_pipeline on WebGPU is a different matter: slot_access cannot check it — that table is derived from the shader signature, so it agrees across a variant that a usage-derived layout would reject — and wgpu exposes no entry count for a BindGroupLayout, so an incompatible hand swap surfaces as a bind group mismatch inside the backend rather than as a set_node_pipeline error.

A site that cannot be specialized without moving its bindings must stay universal. That is a missed optimization, not a bug — but it has to be detected rather than assumed.

Gate the backend, do not tier the feature

The awkward part is that the specializations most worth having are the ones that drop a binding. if (has_tint) { read the tint buffer; blend } is the motivating case, and baking has_tint to zero is exactly what makes that buffer unreachable. On a backend with usage-derived layouts, the predictor would decline promotion precisely where it would have paid off most.

That is a reason to gate WebGPU, not to make specialization a property of "better" backends. Usage-derived layouts are not a WebGPU limitation: the backend passes layout: None and takes wgpu's auto layout. An explicit pipeline layout built from the signature-derived WgpuComputeLayout the backend already computes would be stable under baking, because a layout may declare bindings the shader does not reference. Until that lands, the honest position is that Goldy's WebGPU backend cannot support specialization, not that WebGPU cannot.

So promotion is conditional on a backend capability — GpuBackend::compute_pipeline_layout_follows_signature — rather than on a backend allow-list. Vulkan, DX12, Metal, CUDA, and the mock backend report yes; WebGPU reports no while it uses auto layouts, and flips on by itself once it does not; the CPU backend reports no because its host-callable lowering has no bake macros, so a variant would be the same kernel. The predictor stays backend-agnostic either way, and the capability is internal: specialization is an implementation detail, so it does not belong in RuntimeCapabilities where a program would branch on it. A scheme queries it once, on its first submit.

Off switch

Specialization is on by default. GOLDY_SPECIALIZATION=0 (or false / no / off) disables prediction entirely: no history, no warm compiles, no swaps, and every site stays on the program it was recorded with. The gate follows the convention of GOLDY_DISABLE_CB_REUSE — read through validation_env, with a thread-local override (test_support::SpecializationOverride) so tests that count records across many frames can pin it.

This is separate from the backend capability above. The environment variable is a global kill switch for when prediction is suspected of causing a problem; the capability decides where prediction can be correct at all.

What this rests on

The mechanism depends on runtime properties that are already in place:

  • Params-only dirtiness. Swapping a node's pipeline marks a scheme params-dirty rather than structurally dirty, so only partitions whose baked payload changed re-record. The schedule cache keys on bindings, which a pipeline swap does not touch.
  • Scheme::set_node_pipeline. The swap itself, addressed by a stable node identity.
  • ShaderModule::variant. A new module with merged defines, reusing retained source, search paths, optimization level, and layout checks.
  • Overridable scalar reads. The virtual-main macro described above, which is what makes a param bakeable without the shader author participating.
  • Compile off the device lock. ComputePipeline::new runs Slang outside the backend mutex on Vulkan and DX12, which is what makes an off-thread warm compile possible without stalling submits. Pipeline creation itself still holds the lock, and on Vulkan that is the expensive half (see goldy#175), so a warm compile can still contend with unrelated submits until pipeline creation moves out too.

Shader provenance

A dispatch node holds a ComputePipelineHandle and a ComputePipeline holds a backend handle; neither, on its own, says which program it is. So every ShaderModule keeps its compile inputs — source, search paths, defines, optimization level, layout checks — in a shared ShaderProvenance with a process-unique id, and every ComputePipeline built from the module carries an Arc to it. The runtime can therefore compile a variant of the program a site is running after the caller has dropped the module (ShaderModule::from_provenance), and variants are keyed by provenance id plus baked words.

Ownership matters more than it looks: on Vulkan, destroy_compute_pipeline waits for device idle, so dropping a variant is not a background operation. The scheme holds the variant it has promoted; when a variant is unbound (demotion, a wider promotion, a caller swap) it is parked for two further submits before its Arc is released, by which point no retained command list that bound it is still the one being submitted. The bounded cache usually still holds it after that.

Telemetry

ReplayStats gains specialization_warms, specialization_promotions, and specialization_demotions; Scheme::node_is_specialized(NodeId) answers for one site. Demotions are visible in the stats immediately after the set_node_param that caused them. Each transition also emits a tracing event under the goldy target (debug for warm / promote / demote, warn for a failed compile or a pinned site).

Backend differences

Metal and WebGPU do not retain command lists, so partition-level resubmit counters are not a usable predictor signal there. Scheme cleanliness is: whether a submit found the scheme Clean is tracked on every backend, independent of retention, and that is what the streak counts. The specialization mechanism therefore behaves the same everywhere; only the size of the saving differs.

The overridable macro is emitted by every compute lowering — the native push-constant path, the WebGPU uniform-buffer path, and the CUDA kernel-argument path — each defaulting the macro to its own read expression, so baking works the same way everywhere and the launch layout is unchanged in all three. Graphics stages lower scalars without the indirection; specialization is defined for dispatch sites, so that is deliberate.

Follow-ups

  • Cost model. A dispatch-size (or measured-duration) term so tiny dispatches are never promoted. Needs profiles from a real consumer before the threshold is more than a guess.
  • Runtime-level variant sharing. Many schemes running one shader with the same stable words compile it once each today. Sharing needs an owner that does not cycle with Runtime, e.g. variants that hold a weak device reference and a cache that drains on device drop.
  • WebGPU explicit layouts. Building the pipeline layout from the signature-derived WgpuComputeLayout would let the backend report compute_pipeline_layout_follows_signature and enable prediction there with no predictor changes.
  • Graphics stages. Draw sites with scalar params lower without the macro indirection; specialization is defined for compute dispatch sites only.
  • Pipeline creation off the lock. Warm compiles still take the backend mutex for PSO creation (goldy#175).

Non-goals

  • Per-image or per-command program switching inside a mixed command-stream walk. A single dispatch over a heterogeneous command buffer cannot pick a different program per element; that stays a branch inside the shader.
  • Speculating incorrectly. A miss always runs the universal program. There is no "probably fine" path.
  • Named feature flags as the long-term API. has_tint and its successors do not belong in worker topology.
  • A public specialization API. Callers do not name keys, supply hints, or opt sites in. The only control is the off switch.
  • Caching every combination. The runtime holds one promoted variant per site and a bounded PSO cache, not the cross product of every axis.

Static Shader Bounds Analysis

Status: prototype, opt-in (GOLDY_SHADER_VALIDATION=bounds). Warnings only; a finding never fails a compile. Tracking issue: #290.

Slang rejects constant out-of-bounds array indices at compile time, but a dynamic index such as links[link] is only "checked" at runtime — and on most GPUs that means an undefined read, a hang, or a device loss that is hard to attribute back to a shader line. Goldy can do better for a well-defined class of bugs: prove, conservatively, that every dynamic index into a statically sized array stays inside 0 <= index < length, and report everything it cannot prove with the Slang source location and the call path that reaches it.

The motivating bug

groupshared int links[256];

int link = searchPredecessor(gtid.x);            // may be -1 / -2
int parent = select(link >= 0, links[link], link - 1);

select is an eager intrinsic: both arms are evaluated, so links[link] is read even when link < 0. The explicitly guarded form is safe and must not be reported:

int parent = link - 1;
if (link >= 0) {
    parent = links[link];
}

With Slang 2026.13 the scalar ternary cond ? a : b is lowered to control flow (a branch plus a block parameter) already in the front end, so it behaves like the guarded form; only select(...) produces an eager select instruction. The analysis operates on what the compiler emitted, not on the surface syntax, so it follows whatever Slang decides.

Integration point: Slang IR, in Goldy

The issue asked where in Slang's lowering pipeline such an analysis belongs. The first prototype analyzed the SPIR-V Slang emits; it was then retargeted to Slang's own IR and the SPIR-V implementation removed. The options, revisited with what each representation actually offers:

OptionVerdict
Slang IR pass inside a Slang forkSlang has SSA, dominators and loop analysis, but its IR is not a public API — Goldy consumes libslang.so through the C API only. A fork would have to be rebuilt and re-validated on every Slang bump. Rejected.
Upstream Slang contributionMost attractive long-term home for the generic parts. The prototype here is the evidence for that conversation; see Upstream vs Goldy-owned.
SPIR-V output, analyzed by GoldyAlready SSA with structured control flow and OpAccessChain; NonSemantic.Shader.DebugInfo.100 maps back to Slang source. But the analysis only sees the shader after Slang's target lowering: every generic is monomorphized and every helper inlined (so a finding in a library function is reported once per inlined copy, with no call path), struct fields become anonymous member indices, names survive only as debug hints, and it only exists when the Vulkan feature compiles SPIR-V. The first prototype; superseded.
Slang front-end IR, analyzed by GoldyThe .slang-module container Slang serializes for a translation unit (spSetOutputContainerFormat + spGetContainerCode) is a stable, documented format (docs/design/serialization.md, "fossil") with stable opcode names, and with DebugInfoLevel::Standard every instruction carries a source location. The IR is what the front end produced after semantic checking and SSA construction, before target lowering: functions, generics, witness tables, struct keys, structured ifElse/loop/switch, getElementPtr with the array type intact, and name hints on everything the user named. Target-independent, so one analysis covers every backend. No compiler fork; a pure Rust reader (RIFF + fossil) with no new dependencies. Chosen.

What changed in the analysis when the input changed

Slang IR and SPIR-V expose different information, and the analysis was revisited accordingly rather than ported one-to-one:

  • Calls are real. Nothing is inlined in the front-end IR, so an intraprocedural analysis would know nothing about searchPredecessor(...). The analysis became interprocedural and context-sensitive: each callee is analyzed with the argument intervals of the call site, summaries (return value plus what was written through out/inout parameters) are memoized per (function, argument intervals, generic substitution), and a finding in a helper is reported once, with the union of the ranges over every calling context and the shortest call path from the entry point. Recursion, depth and the number of contexts per function are bounded.
  • Generics are still generic. workgroup_reduce<T : IMonoid, let N : int> is one body; a call is specialize(generic, uint, 64, witness). The analysis carries the substitution into the body: N folds to a constant (also inside constexpr* length expressions such as 1 << LG_N), values of type T take the argument's integer shape (or struct layout), and T.combine(...) is a lookupWitness(table, key) resolved through the witness table to the concrete conformance — including one defined in the shader for its own struct.
  • Imported modules are separate containers. Slang does not embed goldy_exp in the translation unit's IR; it references its declarations through import-decorated stubs carrying the mangled name. Goldy compiles each imported module (transitively, discovered from those names) to its own container, caches it, and links everything into one instruction space by redirecting each stub to the export-decorated definition. After linking, cross-module calls, generic arguments, witness tables and struct field keys are ordinary intra-module references. Unresolvable stubs stay opaque and are named in the diagnostic (no body available).
  • Locals are memory, not phis. The front end keeps var/store/load for constructors and out parameters, and copies by-value aggregate parameters into a local. Non-escaping locals are read through their stores (field-precise, so struct-valued values keep per-field ranges); escaping ones are unknown.
  • No common-subexpression elimination. i + 1 in a guard and i + 1 in the index are distinct instructions. A value-numbering pass canonicalizes structurally identical pure instructions so that dominating facts apply to every copy of the expression they mention.
  • Debug info is in-band. With debug info Slang mirrors stores into by-value aggregate parameters onto DebugVar pseudo-variables; those getElementPtrs are not real accesses and are skipped, otherwise every such site would be counted twice.
  • Built-ins are still seeded from the module ([numthreads] for SV_GroupThreadID / SV_GroupIndex, [1, 128] for WaveGetLaneCount(), the type range for everything else) and core-module intrinsics (min, clamp, abs, countbits, firstbitlow, wave queries, ...) are modeled by name; firstbitlow(N) on a constant is exact, which is how generic library code spells log2(N) in a loop bound.

The analysis compiles the shader a second time as a separate Slang request with debug info and container output; the production bytecode handed to the driver is untouched. Because the front-end IR precedes optimization, the optimization level and target of the production compile do not affect the analysis; it runs for every target.

Analysis model

The code is split into a rule-agnostic IR toolkit and the rules that use it:

  • src/slang/ir/ — the reader and everything a static check over Slang IR needs regardless of what it checks: riff.rs / fossil.rs (container and serialization format), module.rs (instruction tree, decorations, name hints, types with generic substitution, linking a translation unit with its imports into one instruction space), source_loc.rs (Sdeb debug chunks), cfg.rs (per-function CFG, block-parameter incomings, dominators).
  • src/slang/shader_validation/ — the checks. mod.rs parses GOLDY_SHADER_VALIDATION into a ShaderChecks set and runs the selected rules over a linked module (validate); bounds/ is the rule described here (bounds/analysis.rs holds the interval domain and the analysis).
  • SlangCompiler::validate_shader drives the compiles (the translation unit and, transitively, the modules it imports) and SlangCompiler logs the findings once per distinct compile.

Adding another check means another rule module under shader_validation/, a flag in ShaderChecks, and a field in ShaderValidationReport; the IR toolkit does not change.

  1. Values. Every integer scalar/vector gets an interval per lane; structs are tracked field by field; everything else is opaque. Entry-point parameters are seeded from their system-value semantic and [numthreads].
  2. Interval propagation. A flow-insensitive ascending fixpoint with widening followed by a narrowing pass over each function. Block parameters (Slang's phis) join their incoming values, each evaluated under the facts that hold on its edge, so the back edge of for (i = 0; i < 8; i++) contributes [1, 8] rather than a wrapped increment and verts[i] is proven. Arithmetic that may wrap collapses to the type range (sound, not precise).
  3. Path-sensitive refinement. For each dynamic index the dominator tree is walked for ifElse / conditionalBranch / switch conditions that dominate the access (index >= 0, index < N, a >= b, the bool-phi lowering of && / ||, ...) and the index expression is re-evaluated under them, including the relational rule a >= b ⇒ a - b ∈ [0, hi(a) - lo(b)] that workgroup scans rely on (scratch[i - stride] under if (i >= stride)).
  4. Interprocedural. Calls to functions with bodies (in this module or a linked one, generics and witness-table dispatch included) are analyzed in the calling context as described above.
  5. Check. Every getElementPtr / getElement index into a fixed-length array, vector or matrix must satisfy 0 <= index <= length - 1. Constant indices are skipped (Slang already validates them). Runtime arrays (buffers) are not checked: their length is a runtime property of the binding.

Every finding carries the array name (variable plus struct member path; a by-value array parameter is named after the parameter), the array length, the proven index range when it is narrower than the type range, the Slang source location (through Goldy's #line mapping for virtual_main wrappers), the call path, and a provenance note naming what the index depends on that the analysis cannot bound (SV_VertexID, WaveGetLaneCount(), a buffer load, groupshared memory, a float-to-int conversion, a function without a body, a widened loop-carried value). Provenance follows arguments back through the caller, so an index that is a field of a ThreadId parameter is attributed to the SV_DispatchThreadID the wrapper built it from.

What the prototype proves and reports

Covered by src/slang/shader_validation/bounds/tests.rs (every case compiles real Slang):

PatternResult
select(link >= 0, links[link], link - 1) with searchPredecessor as a real callreported, range [-2, 253] (may be negative)
if (link >= 0) parent = links[link];proven
cond ? links[link] : x (scalar ternary)proven — lowered to control flow
links[max(link, 0)], links[i & 255], links[min(i, 255)], links[i % 256]proven
if (i < 256) alone on a signed index from a bufferreported (lower bound missing)
if (i >= 0 && i < 64) on a signed index from a bufferproven (bool-phi conjunction)
Workgroup inclusive scan if (lid >= s) x += scratch[lid - s], reduce, XOR butterflyproven
Same scan with the lid >= s guard removedreported
Padded dispatch: if (gid >= n) return; then scratch[gtid.x]proven
Padded dispatch: sentinel uint idx = gid < n ? gid : 0xffffffff; scratch[idx]reported, depends on SV_DispatchThreadID
for (int i = 0; i < 8; i++) verts[i] (function-local array)proven
static const float2 quad[6]; quad[vertex_id]reported, depends on SV_VertexID; quad[vertex_id % 6] proven
palette[int(hue * 6.0)]reported, depends on a float-to-int conversion
totals[gtid.x / WaveGetLaneCount()] with totals[32], 256 threadsreported, [0, 255] — safe only if the subgroup has ≥ 8 lanes
Helper called from two sites, sh[i] for i ∈ [0, 63] and i ∈ [0, 127]reported once, range [0, 127], call path named
Struct argument p.i with a range, indexed in the calleeproven (fields keep their ranges)
void f(out uint i) result used as an indexanalyzed through the summary
wrap_read<64> vs wrap_read<128> (sh[i % N])proven / reported — per specialization
[goldy_compute] wrapper building GroupThreadId through goldy_expgtid.x keeps its [numthreads] bound
workgroup_reduce(gtid.x, gtid.x, scratch) from goldy_exp, T = uintproven (3 sites inside the library body)
Same with a shader-defined struct Pair : IMonoid whose combine indexes a tablereported in Pair.combine, path cs_main -> _goldy_user_cs_main -> workgroup_reduce, depends on groupshared memory
pick<8>(i, uint t[1 << LG_N])length 256 evaluated from the generic argument

Corpus evaluation

All 33 shaders under shaders/ (51 entry points; triangle.slang and the goldy_exp library itself do not compile as stand-alone entry points) were analyzed. Of 44 dynamic indices into fixed-length aggregates, 23 are proven and 21 are reported:

ShaderFindingAssessment
game_of_life_render, particle_render, rain_snow_render, starfield_render (7 sites)static const vertex table indexed by SV_VertexIDUnprovable from the shader: correctness depends on the draw's vertex count. The idiomatic quad[vertex_id % 6] is proven; documented as the recommended form.
starfield_render (2 sites)galaxyColors[int(hue)], galaxyColors[int(hue) + 1]Safe if hash1 < 1.0; the analysis has no float intervals. Real dependency on a float invariant the code does not state.
goldy_exp/collectives.slang workgroup_inclusive_scan_wave_uint_sum (4 sites, reached from test_collectives)scratch_wave_totals[32] indexed by local_ix / lanes and i < 256 / lanesTrue precondition: the function's own comment says sizes <= 32 cover 256 threads only for subgroups of 8+ lanes. Vulkan only guarantees [1, 128]. Encoding the assumption (lanes = max(WaveGetLaneCount(), 8)) would make it provable.
goldy_exp/collectives.slang workgroup_upper_bound (1 site)prefix_sums[probe - 1] in a branchless binary searchFalse positive. probe - 1 ∈ [0, N - 2] holds by a summation invariant over the loop (ix accumulates halving strides) that interval analysis with widening cannot express. Reported with depends on a loop-carried value the analysis could not bound.
test_goldy_exp_interlocked.slang (7 sites)groupshared sh[64] indexed by ThreadId.x (SV_DispatchThreadID)Latent bug in a compile-only test: correct only when exactly one workgroup is dispatched. One source access per line; SPIR-V had merged them into one access chain.

Barrier-heavy algorithms (test_collectives, test_monoids, test_algebra: workgroup_reduce, workgroup_inclusive_scan, mapped_workgroup_reduce, workgroup_broadcast, prefix sums — 20 dynamic indices inside goldy_exp generics, all reached through import goldy_exp and specialized per call) account for the 5 library findings above; the other 15 are proven. Every finding in a library body is reported once with the call path, where the SPIR-V analysis had reported each inlined copy.

The findings agree with the SPIR-V prototype's site for site; the counts differ only because the front-end IR has one access per source expression (no CSE) and one body per generic (no inlining). Analysis time is 60–150 ms per entry point including the extra compiles.

Re-run with cargo test --lib shader_validation::bounds::tests::survey -- --ignored --nocapture.

Known limitations

  • No float intervals. Any index derived from a float conversion is unknown.
  • Memory model. Only non-escaping locals (and pointer parameters) are read through their stores; groupshared and buffer contents are unknown (only their indices are checked), and an address that escapes to an intrinsic makes its variable unknown.
  • Intervals only. No relations between variables beyond dominating comparisons, no loop summation invariants, no cross-thread reasoning.
  • Bounded interprocedural analysis. Call depth, contexts per function and total function analyses are capped; past the caps a callee is analyzed once with unconstrained parameters. Recursion is cut at the first repeated function. Functions whose bodies are unavailable (target intrinsics not modeled by name, modules whose source is not on a search path) are unknown and named in the diagnostic.
  • Generic witness tables that are themselves generic (extension<T> ...) are not followed; the dispatched method is unknown.
  • Subgroup size. WaveGetLaneCount() is [1, 128] per Vulkan. Shaders that assume a minimum lane count should state it (max(lanes, 8)).
  • Slang serialization format. The reader follows Slang 2026.13's fossil layout and stable opcode names (stable_names.rs, generated from slang-ir-insts-stable-names.lua); a Slang bump that changes either fails the analysis with a Malformed error (logged at debug), never the compile.

Upstream vs Goldy-owned

Decision for now: Goldy-owned analysis stage over Slang's serialized IR, with an eye to upstreaming the generic pieces.

  • The pieces that are Goldy-specific and stay here: reading the container, compiling and linking imported modules, mapping locations back through the virtual-main rewrite, the GOLDY_VALIDATION plumbing, seeding built-ins from [numthreads], and the corpus-tuned diagnostics (provenance notes, parameter-source attribution, naming through by-value copies).
  • The pieces that are generic and would serve every Slang user: interval propagation with edge-refined phis and narrowing, dominating-comparison refinement including the bool-phi &&/|| rule and the a >= b ⇒ a - b >= 0 relational rule, context-sensitive summaries over specialize / lookupWitness, and the getElementPtr check with a "cannot prove 0 <= index < length" warning. Working on Slang IR now means the analysis is expressed in the terms an upstream pass would use (Slang has SSA, dominators, IRIntegerRelation and loop analysis of its own), so the false-positive profile above can be discussed with the Slang maintainers on equal footing.
  • What we would need from upstream regardless: a way to surface the analysis through the C API so Goldy can keep formatting locations through its own #line mapping.

Running it

Static shader checks have their own variable, GOLDY_SHADER_VALIDATION, and are deliberately not part of GOLDY_VALIDATION=all: the runtime categories there are cheap invariant checks on Goldy's own state, whereas this is a second compile plus a whole-program analysis per shader whose findings mean "not proven", not "wrong". The value is a list of check names processed left to right: all, bounds, -bounds (exclude), none.

# Warn on every dynamic index the analysis cannot prove in bounds
GOLDY_SHADER_VALIDATION=bounds cargo run --features examples --example metaballs
RUST_LOG=goldy::slang=debug GOLDY_SHADER_VALIDATION=all cargo test

# Runtime validation and static checks are independent
GOLDY_VALIDATION=all GOLDY_SHADER_VALIDATION=all cargo test

# Keep the IR container the checks looked at
GOLDY_DUMP_SHADERS=/tmp/shaders GOLDY_SHADER_VALIDATION=bounds cargo run ...   # {entry}_ir.slang-module

# Print the linked IR as text / re-run the corpus survey
GOLDY_IR_DUMP=shaders/test_collectives.slang cargo test --lib slang::ir::tests::dump_ir -- --ignored --nocapture
cargo test --lib shader_validation::bounds::tests::survey -- --ignored --nocapture

Findings are logged once per distinct compile (including shader disk-cache hits) at warn under the goldy::slang target, tagged with the check:

shader validation (bounds): possible out-of-bounds index into `links[256]`: index range [-2, 253] (may be negative) (depends on SV_GroupThreadID) at shader.slang:11:1 in `cs_main`
shader validation (bounds): possible out-of-bounds index into `scratch_wave_totals[32]`: index range [0, 255] (may exceed 31) (depends on the result of `WaveGetLaneCount()`) at shaders/goldy_exp/collectives.slang:123:28 in `workgroup_inclusive_scan_wave_uint_sum` (called from cs_main -> _goldy_user_cs_main)

Programmatic access: SlangCompiler::validate_shader(.., ShaderChecks) compiles a shader (and the modules it imports) and returns a ShaderValidationReport; its bounds field is the BoundsReport with checked_accesses, proven_safe and the list of BoundsDiagnostics. goldy::slang::shader_validation::validate runs the checks on .slang-module bytes produced elsewhere.

Goldy vs wgpu

Both Goldy and wgpu are Rust GPU libraries with multi-backend support. They make different tradeoffs that suit different use cases.

At a Glance

wgpuGoldy
IdentityWebGPU implementation for RustModern Rust GPU library
Spec governanceW3C WebGPU specificationIndependent, opinionated
Browser supportYes (WebGPU)No
Minimum hardwareWide compatibility (Vulkan 1.0+)Modern only (Vulkan 1.4+, DX12, Metal 2+)
Shader languageWGSL (primary), SPIR-V, GLSL, nagaSlang (compiles to SPIR-V, DXIL, MSL)
Resource modelDescriptor-based (bind groups)Typed bindless
SynchronizationManual pass orderingRetained scheme (dependency graph)
Metal supportVia MoltenVK or wgpu-halNative Metal backend
Compute modelSupported but secondaryFirst-class (compute-to-surface)

Resource Binding: Descriptors vs Bindless

wgpu uses bind groups — the WebGPU equivalent of Vulkan descriptor sets. You declare a bind group layout, create bind groups that match it, and bind them before each draw or dispatch:

#![allow(unused)]
fn main() {
// wgpu: declare layout, create group, bind before draw
let layout = device.create_bind_group_layout(&desc);
let group = device.create_bind_group(&wgpu::BindGroupDescriptor {
    layout: &layout,
    entries: &[wgpu::BindGroupEntry { binding: 0, resource: buffer.as_entire_binding() }],
    ..
});
pass.set_bind_group(0, &group, &[]);
}

Goldy uses bindless access. Resources get a slot index at creation time, and shaders access them directly by index. There are no layouts, groups, or binding calls:

#![allow(unused)]
fn main() {
// Goldy: bindless parcel already has a slot; bind it in the scheme
let parcel = runtime.acquire_buffer_with_data(&data, BufferKind::Scattered)?;
pass.with_parcel(&parcel, NodeAccess::Read);
}

The bindless approach eliminates an entire layer of API surface and the pipeline layout permutations that come with it.

Synchronization: Manual vs Dependency Graph

wgpu provides implicit synchronization within a render/compute pass but requires you to order passes correctly. Resource transitions between passes are handled by wgpu internally, following WebGPU's implicit rules.

Goldy uses a retained scheme — a dependency graph recorded once and submitted each frame. You declare nodes and their resource dependencies; Goldy derives barriers, layout transitions, and execution order. This gives the runtime a global view of the frame for optimal scheduling and makes synchronization bugs structurally impossible.

Shader Language: WGSL vs Slang

wgpu's primary shader language is WGSL, the WebGPU Shading Language. WGSL is designed for safety and portability across web and native targets, but it lacks features like modules, generics, and automatic differentiation.

Goldy uses Slang exclusively. Slang compiles a single source file to SPIR-V (Vulkan), DXIL (DX12), and MSL (Metal). It provides modules with true separate compilation, generics, and HLSL-familiar syntax. The goldy_exp shader library builds on Slang's module system to provide shared types and utilities:

import goldy_exp;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Particle> particles, ThreadId id) {
    particles[id.x].position += particles[id.x].velocity;
}

Compute as First-Class Citizen

wgpu supports compute shaders, but the API is oriented around render passes. Compute-to-render workflows require manual buffer management and pass ordering.

Goldy treats compute and graphics as peers. Compute-to-surface is a built-in pattern: a compute dispatch writes to a buffer or texture, and a subsequent render pass reads from it, with the scheme handling the dependency automatically.

Metal: Native vs MoltenVK

wgpu supports Metal through its wgpu-hal Metal backend or via MoltenVK (Vulkan-on-Metal translation). MoltenVK adds a translation layer that can introduce overhead and compatibility limitations.

Goldy has a native Metal backend that uses Metal APIs directly — Argument Buffers Tier 2 for bindless, MSL compiled from Slang, and native Metal types throughout. No translation layer sits between Goldy and the Metal driver.

Architecture

wgpu:

Application → wgpu (WebGPU API) → wgpu-hal → Vulkan / Metal / DX12 / WebGPU

Goldy:

Application → Goldy (native API) → Vulkan 1.4+ / Metal 2+ / DX12 (shipped); CUDA / WebGPU (in progress); Tenstorrent (planned)

wgpu implements the WebGPU specification faithfully, then maps it onto each backend through an internal HAL. Goldy talks to each backend directly using native idioms.

When to Choose Which

Choose wgpu when:

  • You need browser deployment via WebGPU
  • You need to support older GPUs or wide device compatibility
  • You want the stability of a specification-driven API
  • You need the wgpu ecosystem (examples, community, tooling)

Choose Goldy when:

  • You target only modern desktop/mobile hardware (2018+)
  • You want a minimal API surface with bindless as the default
  • You want native Metal without a translation layer
  • You want Slang's module system and shader language features
  • Compute workloads are central to your application

Both libraries are valid choices — the right one depends on your hardware requirements, deployment targets, and whether you value broad compatibility or API simplicity.

Target Hardware

Goldy targets modern GPUs exclusively. This is a deliberate design choice — by requiring hardware from roughly 2018 onward, Goldy can use bindless descriptors, dynamic rendering, and coherent caches as baseline assumptions rather than optional features.

Backend Requirements

Vulkan 1.4+

Goldy requires Vulkan 1.4, which promotes several extensions that were optional in earlier versions to core:

FeatureVulkan historyGoldy usage
Dynamic renderingVK_KHR_dynamic_rendering (1.3)No render pass objects
Descriptor indexingVK_EXT_descriptor_indexing (1.2)Bindless resource access
Buffer device addressVK_KHR_buffer_device_address (1.2)64-bit GPU pointers
Synchronization2VK_KHR_synchronization2 (1.3)Simplified barrier model
Push descriptorsCore in 1.4Efficient uniform updates

Supported hardware:

  • NVIDIA: Turing and later (RTX 2000 / GTX 1600 series, 2018+)
  • AMD: RDNA 1 and later (RX 5000 series, 2019+)
  • Intel: Xe architecture and later (Arc, 2022+)
  • Qualcomm: Adreno 650+ (2019+, driver dependent)

DX12

Goldy's DX12 backend requires:

RequirementDetails
D3D12 Enhanced BarriersWindows 11 + WDDM 3.0+ driver
ResourceDescriptorHeapSM 6.6 bindless (Shader Model 6.6)
Root constantsPush constants equivalent

Enhanced Barriers are mandatory — Goldy does not fall back to legacy resource state transitions. This effectively requires Windows 11 with a modern driver.

For software rendering and CI, Goldy supports the WARP software rasterizer via GOLDY_DX12_FORCE_WARP=1.

Metal Tier 2+

Goldy's Metal backend is native (no MoltenVK) and requires Argument Buffers Tier 2 for bindless resource access:

RequirementDetails
Argument Buffers Tier 2Bindless via ParameterBlock
MSL (via Slang)Slang compiles directly to Metal Shading Language

Supported hardware:

  • Apple Silicon: All models (M1/M2/M3/M4, A14+)
  • Intel Macs: 2017+ (different iGPUs; some very early Intel UHD may not qualify)
  • AMD discrete GPUs in Macs: 2015+

Older Intel integrated GPUs (pre-2017 Macs) are not supported — they lack Argument Buffers Tier 2.

What "Modern GPU" Means for Goldy

Goldy's hardware floor is defined by a set of architectural capabilities, not specific product names:

CapabilityWhy Goldy needs it
Coherent L2 cacheNo manual cache flush/invalidate logic
Bindless descriptorsSingle global descriptor model, no set layouts
Dynamic renderingNo render pass objects or framebuffer compatibility
64-bit buffer addressesDirect pointer access in shaders
Unified or REBAR memorySimplified CPU-GPU data transfer

GPUs from roughly 2018 onward universally support these features. The specific API version requirements (Vulkan 1.4, DX12 Enhanced Barriers, Metal Tier 2) are the mechanism by which Goldy enforces this floor.

Optional: ray tracing and mesh shaders

Bindless, dynamic rendering, and buffer device addresses are required. Hardware ray tracing and mesh shaders are not — GTX 16-series, RDNA1, and many iGPUs meet Goldy's floor without RT cores or mesh shaders.

Query them on Adapter::capabilities() / Runtime::capabilities() (RuntimeCapabilities):

FlagMeaning
ray_queryInline ray queries (Vulkan VK_KHR_ray_query, DXR 1.1, Metal supportsRaytracing)
ray_tracing_pipelinesRT pipelines (Vulkan VK_KHR_ray_tracing_pipeline, DXR 1.0+). Metal reports false — use ray_query.
mesh_shadersMesh shaders (Vulkan VK_EXT_mesh_shader, DX12 mesh tier 1, Metal Apple7 / Mac2 / Metal3)
amplification_shadersTask / amplification / object shaders with mesh

These bits report adapter hardware. Inline RayQuery ([goldy_compute]) records on Vulkan, DX12, and Metal when ray_query is set. Ray-tracing pipelines (RayTracingPipeline + Scheme::trace_rays) record on Vulkan and DX12. Mesh draws (MeshPipeline + dispatch_mesh) record on Vulkan, DX12, and Metal when mesh_shaders is set.

Additional Backends

BackendStatusNotes
CUDAIn progressNVIDIA compute prototype; cuda Cargo feature
WebGPUIn progressCross-platform prototype via wgpu; webgpu Cargo feature
TenstorrentPlannedTorus Fondaco runtime design (not yet implemented)

These backends are not yet supported for production use. See Backend Architecture for details.

What This Excludes

ExcludedReason
NVIDIA GTX 900 series (Maxwell)No Vulkan 1.4 support
AMD GCN (RX 400/500)Driver support ended; limited bindless
Intel Gen9 (HD 500/600)Incomplete Vulkan feature coverage
Intel integrated GPUs pre-2017 (Mac)No Argument Buffers Tier 2
Pre-Windows 11 DX12No Enhanced Barriers

Checking Compatibility

Goldy reports unsupported devices at initialization:

#![allow(unused)]
fn main() {
let instance = Instance::new()?;

for adapter in instance.enumerate_adapters() {
    println!("{}: {:?}", adapter.name, adapter.device_type);
}

// request_runtime returns an error on unsupported hardware
let device = instance
    .request_adapter(&RequestAdapterOptions::default())?
    .request_runtime(&RuntimeDescriptor::default())?;
}

The Tradeoff

By drawing a line at modern hardware, Goldy avoids the fallback paths, compatibility checks, and feature-level negotiation that dominate traditional GPU libraries. Every code path in Goldy assumes the full feature set is available. This keeps the implementation small and the API surface predictable.

The cost is clear: Goldy cannot run on the long tail of older hardware. For applications that need broad device support, wgpu is the better choice.

Fondaco Machine (research)

Goldy realizes the Fondaco Machine — an abstract model of cooperative GPU computation. These chapters are the normative research material adapted for this book.

Terminology

Vocabulary and status labels used throughout the Fondaco chapters and the rest of this book.

AuthorityDocument
Machine semanticsMachine Specification
Goldy realizationGoldy Runtime Mapping
Shipped behaviorGoldy source, tests, and examples

Status labels

LabelMeaning
ShippedAvailable in the public Goldy crate today (0.2.x)
DesignedSpecified and intended; not yet implemented, or only partially implemented
ExperimentalBehind a feature flag, alpha binding, or unstable API
SpeculativeResearch or exploration; not committed to the roadmap
HistoricalSuperseded design kept for context

Do not describe Designed, Experimental, or Speculative capabilities as if they were Shipped.

Fondaco machine terms

TermDefinition
MerchantThe sovereign client: describes parcels and executes schemes; owns every parcel
SchemeFirst-class computation: dispatches plus precedences (a partial order)
DispatchAtomic unit of work admitted to the machine
ScriptOpaque procedure evaluated by a computing dispatch
Yielding scriptScript with structured yield points where the runtime may be petitioned
ParcelStable identity for data held by the runtime in trust for the merchant
Ownership / claimAccess right over a parcel for the duration of a dispatch (public, private, or private-inaugural)
LedgerRuntime standing record of claims across schemes (not merchant-addressable)
GateInterval between adjacent dispatches where the runtime has full intervention powers
ExchangeStable mediated relationship with a foreign subsystem; each execution may publish a linear claim
WarehouseRuntime-imposed bound on the total extent of parcels the merchant may hold
PetitionStructured service request filed at a yield point

Goldy API map

Fondaco termGoldy type / conceptNotes
SchemeScheme, internal GraphIRPublic recording and submission API
DispatchScheme node (compute, render, copy, clear, present)Workgroup grid for compute/render
ScriptSlang via [goldy_*] virtual entry pointsSole script language (Goldy choice)
ParcelParcel, Buffer, TextureStable handles; bindless indexing is backend-internal
OwnershipNodeAccess on scheme nodesPrecedences derived from access modes
LedgerCross-submission sync (ParcelStamp, timeline)Crate-private; clients use settlement APIs
GateSubmission gate, Context::boundary_crossedEpoch-driven reclamation
ExchangeSurfaceExchange, MemoryExchangePresent: (&mut submission >> &transaction).take()? (canonical: Transaction → Claim → consume / discard). Deposit: (&deposit << &data)? tenders this submission; internal Claim consumed at copy dispatch
Host claimPendingHostRead, HostView(&mut submission >> &parcel).take::<T>() — mapped read or staged copy; public CPU ownership between gates
WarehouseBudgetPolicy, VramAllocatorBound on committed parcel extent
LeaseLease<LeaseTexture>, Lease<LeaseBuffer>, Lease<LeaseRenderTarget>, PresentLeaseTemporary tenancy minted by the lessor (Context / surface pool); schemes intern on first use

Internal terms

TermLocationRole
GraphIRtask_graphInternal scheme representation
Wave / partition analysistask_graph::analysisSubmission partitioning and transient coloring
Bindless heapBackendsDescriptor indexing; not public ABI

Reading order

  1. Machine Specification — normative semantics
  2. Goldy Runtime Mapping — what Goldy ships vs designs
  3. Render Passes and Schemes — raster grain vs pass clustering
  4. Design Thesis — why this model on modern GPUs
  5. The rest of this book — tutorials, programming model, and APIs

Machine Specification

Status: Draft v0.12

This chapter specifies the Fondaco abstract machine: what a merchant is, what parcels and ownership mean, what execution is, and what a runtime must do. It does not specify a particular hardware mapping, calling convention, or API. Those belong to implementations — Goldy's mapping is in Goldy Runtime Mapping.

The machine is a positive specification: it states what is, and what is not stated does not exist. Analogies to other abstract machines appear only in the appendix and are leaky. The Fondaco terms in the body are authoritative. See also Terminology.

1. The machine

The Fondaco machine is an abstract machine for cooperative computation, implemented by a runtime.

Parcels are data. Schemes are computation.

A merchant describes parcels and executes schemes via dispatches. The merchant is the sovereign — every parcel belongs to it; the runtime holds them in trust. The merchant assigns ownership claims through schemes; the runtime manages their physical realization.

The remainder of this chapter specifies schemes, dispatches, scripts (including yielding), parcels and ownership, exchanges, gates, petitions, and the latitude runtimes have to transform schemes.

2. The scheme

A scheme consists of:

  • A set of dispatches
  • A set of precedences between dispatches

A precedence A → B asserts that dispatch A must complete before dispatch B begins. The transitive closure is a partial order over the dispatches.

Two dispatches are unordered if neither is downstream of the other. The runtime may execute unordered dispatches in any order, including concurrently or fused.

A scheme is first-class data. Schemes compose: a scheme may contain another as a sub-scheme, and the composition of two schemes is a scheme.

A scheme is well-formed if its precedence set is consistent with its ownership claims (§5). Specifically: no two unordered dispatches may claim the same parcel where at least one holds private ownership. A merchant may add precedences beyond those ownership requires — for throughput or occupancy — but may not omit a required precedence.

Presenting an ill-formed scheme has unspecified behavior. Conforming runtimes may refuse to admit it (§10). Goldy refuses ill-formed schemes rather than producing unspecified results.

The precedence set need not be acyclic. A cyclic scheme describes a non-halting computation; whether to admit it is left to the runtime.

3. Dispatches

A dispatch is the atomic unit of work admitted to the machine. Once begun, it runs to completion. It is not preempted or reordered internally.

A dispatch runs over parcels. If arbitrary computation is required, the runtime schedules a script (the execution model is left unspecified); otherwise the runtime performs the dispatch directly.

4. Scripts and yielding

A script is the procedure a computing dispatch evaluates. The language is unspecified by the machine; runtimes may fix one or more. The script is opaque to the runtime: control flow is not visible. A script may keep private resources that are invisible to the runtime during a dispatch.

Yielding

A script may yield if it contains yield points that transfer control to the runtime and later resume.

At a yield point:

  1. The dispatch reaches the yield point and petitions the runtime (§7).
  2. The script's private resources are undisturbed while suspended.
  3. The runtime may modify claims while suspended, as long as it restores the state the dispatch observed before the yield.
  4. The runtime unsuspends the dispatch; the script resumes from the yield point.

A yield point does not create a gate. The dispatch has not ended. Visibility is limited to servicing that petition and scheduling — not the full gate powers of §6.

Non-yielding scripts

A script with no yield points runs start to finish with no runtime intervention. The runtime is invisible for the duration of a non-yielding script.

Goldy today ships only non-yielding scripts; yielding is Designed (see Goldy Runtime Mapping).

5. Parcels and ownership

A parcel is a stable identity for data held by the runtime in trust for the merchant. Parcels are the sole channel of communication between dispatches within a scheme or across schemes. Dispatches are not aware of physical realization; they rely on stable identity.

Each parcel has:

  • A type, fixed at creation (type system unspecified by the machine)
  • A size — extent or shape, possibly hinted, possibly fixed
  • A claim — ownership granted by the merchant

Ownership

Ownership is expressed as claims: access rights for the duration of a dispatch.

  • Public — read. Multiple dispatches may hold public ownership of the same parcel concurrently.
  • Private — read and write. Exclusive: no other concurrent claim on that parcel.
  • Private-inaugural — write without depending on prior contents. Exclusive, but no precedence from a prior owner is required; prior state is abandoned.

Ownership transfer

If dispatch A holds a private claim on parcel X and dispatch B (ordered after A) claims X, ownership transfers at the gate between them. If no later dispatch claims a parcel after its last holder completes, claims drop and ownership reverts to the runtime, which may destroy the parcel if the merchant no longer needs it.

The ledger

The ledger is the runtime's standing record of claims over parcels. Well-formedness (§2) is a property of one scheme in isolation; the ledger is the aggregate account between schemes. When one scheme completes and another is admitted, later claims serialize against prior owners the ledger records.

At every instant, at most one private claim — or any number of public claims — stands over a given parcel. The runtime mutates the ledger only by admitting schemes, acting at gates (§6), and settling exchanges. Pending foreign access from an exchange is a standing constraint until that access expires.

The ledger is an invariant, invisible to the merchant. The machine does not require a particular data structure — only that the runtime behave as though such a record is conserved.

Exchanges

An exchange is a stable, runtime-mediated relationship between a scheme and an entity outside the machine. Through an exchange, a scheme may periodically hand parcels to, or receive parcels from, a foreign subsystem.

Establishing an exchange records the relationship and its ownership constraints; it does not itself perform a foreign operation. Executions may produce exchange claims, distinct from ownership claims.

Settlement is either:

  • Consume — perform the foreign operation defined by the exchange
  • Discard — settle without exercising it

Settlement is terminal even if it reports failure. A runtime must define safe settlement for claims abandoned by the merchant. Representation and delivery of claims are defined by the exchange, not the machine.

Foreign access constrains ownership until the runtime knows the access has expired (reads) or completed with a usable produced state (writes). Which foreign subsystems exist and how completion is observed are unspecified by the machine. A runtime with no exchange conventions simply permits no foreign I/O.

The warehouse

The runtime need not provide an infinite warehouse. It may impose a warehouse — a bound on total parcel extent the merchant may hold. The merchant remains sovereign within that bound.

Exceeding the warehouse is a runtime-defined condition; the runtime need not admit the scheme. The warehouse may expand or contract. On contraction, the runtime may reclaim medium at the next gate for parcels whose claims have been relinquished, but must not destroy a parcel whose claim the merchant still holds.

A preferred warehouse size is a hint: the runtime honors min(declared, available). No preference means the runtime default.

The physical medium

Physical backing is managed entirely by the runtime. At any gate it may reorganize, relocate, or reclaim medium. None of this is observable to the merchant: only claims and parcel identity are preserved across physical activity.

6. Gates

A gate is the interval between two adjacent dispatches in an execution order.

At a gate, the runtime may:

  • Transfer ownership between dispatches
  • Inspect or modify contents of parcels it holds in trust
  • Relocate physical medium
  • Reclaim medium for parcels whose claims have been relinquished
  • Insert additional dispatches into the scheme

Within a dispatch (including at yield points), the runtime may exercise gate powers only if they are not observable to the dispatch upon resumption.

7. Petitions

A petition is how a dispatch requests a service from the runtime. It may be filed only at a yield point (§4). The dispatch suspends, the runtime services the petition, then the dispatch resumes.

At a yield point

  1. The script reaches the yield point and signs the petition.
  2. The runtime may perform the service — including scheduling other work, delivering parcels to the merchant, or internal bookkeeping.
  3. The runtime resumes the script with script-state intact.

Petition conventions (encodings, services offered, how yield points are declared) are unspecified by the machine. A runtime that defines none provides no services beyond execution.

8. Scheme transformations

The runtime may transform a scheme before or during execution if observable behavior is unchanged. For any well-formed scheme it admits, it may:

  • Fuse adjacent dispatches into one, producing the same parcel states
  • Split one dispatch into several with precedences, producing the same parcel states
  • Reorder unordered dispatches freely, including concurrent execution
  • Elide dispatches whose effects can be derived without execution
  • Insert bookkeeping dispatches that do not alter merchant-observable parcel states

These are algorithms over schemes (§9). Merchants express natural granularity; the runtime reshapes for hardware. Final parcel states are the invariant.

9. Algorithms over schemes

A scheme is first-class data. Execution — producing parcel states consistent with the partial order — is the defining algorithm, but not the only one.

Others include (without limit): fusion, splitting, specialization, differentiation, distribution across runtimes, scheduling annotations, analysis without execution, and composition.

A runtime need only provide execution; the machine admits all possible algorithms over schemes.

10. Conformance

A runtime conforms if and only if, for every well-formed scheme it admits, it produces parcel states consistent with at least one execution that respects:

  • The scheme's partial order
  • Ownership rules in §5
  • Exchange constraints in §5
  • Parcel-identity contracts in §5 and §6
  • Yield-point petition constraints in §7 (no full gate powers at yield points)

A conforming runtime may decline to admit a scheme. It need not support every script language, parcel type, or physical extent — only those it advertises. It need not support yielding scripts.

The machine does not specify performance, latency, energy, or resource consumption — only the meaning of what the merchant executes.

Appendix: Glossary (non-normative)

Loose analogues for readers familiar with other models. The Fondaco terms above are authoritative.

Fondaco termLoose analogueNote
MerchantProgram, application, clientSovereign owner of all parcels
SchemeCommand buffer, render graph, dataflow graphFirst-class; algorithms operate over it
DispatchKernel launch, shader invocationAtomic; runs to completion
ScriptKernel body, shader sourceOpaque except at yield points
Yielding scriptCoroutine, fiber bodyStructured suspend/resume
Yield pointSuspension point, syscall boundaryCollectively transfers control to the runtime
ClaimsDescriptor set, root signature, argument bufferNames mapped to parcels with ownership
ParcelBuffer, texture, resourceStable identity; medium may move
LedgerResource-state / hazard trackerConserved across schemes; not merchant-addressable
ExchangeSwapchain present, DMA to host-owned memoryLinear per-execution settlement
Host claimMapped read / readbackPublic claim by the host between gates
Ownership (public)SRV, sampled image, read-only bindingConcurrent reads
Ownership (private)UAV, storage image, read-write bindingExclusive tenant
Ownership (private-inaugural)Discard/clear load opExclusive write; prior state abandoned
WarehouseRuntime memory budgetBound relative to other merchants
GateFence, barrier, semaphoreFull intervention between dispatches
PetitionSystem call, trapMid-dispatch; limited service, not full gate powers
Scheme transformationCompiler pass, kernel fusionObservable parcel states must be preserved
RuntimeOS kernel, driver, command queueConceptually one agent; may be distributed

Analogues are not equivalences. In particular: a scheme is the program, not merely a scheduling artifact; parcel identity must stay opaque (descriptor heaps are backend-private); petitions at yield points do not grant full gate powers.

Goldy Runtime Mapping

Status: Implementation note for Goldy 0.3.x. Claims use the labels in Terminology.

How Goldy realizes the Fondaco machine from Machine Specification. Goldy is a runtime, not the machine. Where this chapter disagrees with the spec, the spec governs.

Hardware terms (GPU, Vulkan, Metal, DX12, shader, fence) appear because Goldy must speak them. They have no normative meaning in the Fondaco machine.

For why Goldy looks this way, see Design Thesis. For day-to-day usage, start at the Introduction.

1. Goldy realizes a Fondaco machine

Shipped. Goldy is a Rust library that admits dispatches, honors scheme partial orders, maintains parcel identity across physical activity, and acts at gates as the machine requires.

It targets 2020-era heterogeneous compute: a host processor plus one or more GPUs via Vulkan 1.4+, DX12, or native Metal (macOS).

The spec's runtime is a single agent. Goldy exposes that agent as a cloneable Runtime: warehouse owner for retained parcels, shader and pipeline factory, capabilities, and diagnostics. Context is a submission timeline (transient/deposit pools, scheme recording). Backend device handles stay private.

Goldy still splits execution across:

  • Host (Rust): parcel identity, schemes, ledger analysis, gates, exchanges
  • GPU queue: executes admitted dispatches

That split is a substrate artifact, not a machine requirement.

2. Status overview

Machine conceptGoldy realizationStatus
SchemeScheme + internal GraphIRShipped
DispatchCompute / render / copy / clear / present nodesShipped
CPU dispatchScheme::cpu_node — serial host function over staged parcelsShipped (correctness first; staging always copies)
ScriptSlang via [goldy_*] virtual entry pointsShipped
ParcelParcel, Buffer, Texture (stable handles)Shipped
Ownership / claimsNodeAccess → derived precedencesShipped
LedgerCross-submission sync (ParcelStamp, timeline)Shipped (internal)
GateSubmission gate, Context::boundary_crossedShipped
ExchangeSurfaceExchange, MemoryExchangeShipped
Exchange claimPresent: (&mut submission >> &transaction).take()?; deposit: (&deposit << &data)? tenders this submission (internal claim at submit)Shipped
Warehouse / budgetBudgetPolicy, VramAllocatorShipped (partial)
Growable buffersBuffer::resize_to, stable handlesShipped
Retained resubmitClean schemes replay with zero re-recordShipped
Compute-to-surfaceSurfaceExchange::bind_destinationShipped
Pipelined framesFrameOrchestrator, surface depthShipped
Yielding scripts / $yieldSlang intrinsic + petition servicingDesigned
Sub-scheme inclusionScheme::include / Scheme::group — snapshot copy of a child descriptionShipped
Scheme fusion (mega-kernel)Merge adjacent dispatchesDesigned
Raster pass as fused drawsOne RenderPass node per framebuffer epochShipped (finer per-draw nodes: Designed)
Scheme splitting (wavefront)Split at yield pointsDesigned
DefragmentationVramAllocator::defragmentDesigned
Memory-pressure eventsMemoryPressureEventDesigned
Promise / continuation APIIndirect continuation dispatchDesigned
WASI host (goldy-host)GPU to WASM guestsSpeculative
Pre-2020 bindless-free backendTraditional binding backendSpeculative

3. Dispatches

Shipped. A Goldy dispatch is a compute or graphics submission to the accelerator.

  • Script: Slang compiled through virtual_main to SPIR-V, DXIL, or Metal IR. Goldy fixes Slang; the machine does not require it. See Virtual Entry Points.
  • Execution: Workgroup grid (threadgroups on Metal). dispatch_indirect where the backend allows.
  • Claims: NodeAccess on scheme nodes — read, write, read-write — mapped to public / private / private-inaugural ownership.

Non-computing dispatches also Shipped: buffer copy, buffer write, texture upload, buffer clear, present / copy-to-swapchain.

Raster: fused dispatches, not machine-opaque scripts

Shipped. A Goldy render-pass node is one dispatch covering a framebuffer epoch (vkCmdBeginRendering … EndRendering, Metal render encoder, DX12 render pass / RTV span) plus every draw / dispatch_mesh recorded in that epoch.

In the machine, each of those draws may be a dispatch. Goldy fuses them at record time because the 2026 portable ABIs do not expose a gate between draws without checking the attachment back in.

CPU dispatches (Shipped, 0.2.x): Scheme::cpu_node admits a serial host function whose parameter list is the virtual main (&[T] / &mut [T] per bound parcel, then scalars). The machine does not distinguish where a dispatch executes; Goldy realizes host execution by staging bound parcels through readback/upload copies around a fence wait, so the node is a full drain of the device pipeline. See CPU Dispatches.

Goldy does not preserve shader invocation identity across dispatch gates; logical threads must persist state in parcels.

4. Parcels and bindless internals

Shipped. A Goldy parcel is a stable handle. Programs never author raw (category, index) bindless slots.

Bindless descriptor indexing is backend-internal:

  • Rust: Public types are Parcel, Buffer, Texture, Scheme, exchanges. Bindless resolution happens at scheme record / submit.
  • Slang: Typed parameters (Scattered<T>, BufRO<T>, …). virtual_main generates slot packing.

Program-visible are access-pattern categories (Layer B): Scattered, BufRO, Broadcast, Interpolated, DirectSpatial, Filter. See Parcels and Design Thesis.

Parcel identity and reslot

Shipped. Identity is the handle, not the descriptor slot. Handles stay stable across physical growth (Buffer::resize_to), transient aliasing within an epoch, and backend pool rotation.

When backing changes (reslot), Goldy:

  1. Keeps the handle unchanged
  2. Gives the new allocation a new descriptor slot; old slots remain valid until in-flight work retires
  3. Invalidates retained command buffers that embedded stale slots

This follows DX12 / Vulkan / Metal descriptor versioning. See Buffers and VRAM Allocator.

5. Warehouse and memory

Shipped (partial). Physical medium is managed by Runtime (retained warehouse), VramAllocator, and TransientPool. See Runtime-Owned Memory and Transient Allocation.

Goldy distinguishes three quantities:

QuantityOwnerMeaning
Logical warehouseProgram + runtimeSum of parcel extents (Fondaco warehouse)
CommittedRuntimeBytes handed out (commit charge)
ResidentOSBytes in the fast tier now

Budget enforcement keys on committed. Resident enters reactively via OS memory-pressure signals.

Shipped residency models per backend: ManagedAllocation (discrete Vulkan / DX12), PageOnFault (Apple Metal), plus capability queries for resize cost (Constant, PageBind, Copy).

Designed: defragmentation, proactive memory-pressure petition delivery at gates.

6. Exchanges

Shipped. Primary exchange: surface presentation.

#![allow(unused)]
fn main() {
let transaction = surface_exchange.bind_render_target(&mut scheme, &scene_rt)?;
let mut submission = scheme.submit()?;
(&mut submission >> &transaction).take()?; // present
}
  • Binding does not acquire a drawable; acquire runs at submit when the partition needs it
  • (&mut submission >> &transaction).take() is sugar for transaction.claim(&mut submission)?.consume(); the &mut borrow leaves other claims untouched
  • Claim::consume / Claim::discard remain the canonical settlement verbs
  • The program never passes raw GPU addresses to the compositor

Shipped CPU readback: host claims via (&mut submission >> &parcel).take::<T>() (PendingHostRead / HostView). See Settlement and Compute to Surface.

Shipped CPU upload: MemoryExchange::bind_deposit records copy topology once. (&deposit << &data)? (or DepositTransaction::write) prepares an occurrence for this submission; Scheme::submit claims it internally and graph execution consumes the claim at the deposit copy dispatch. Staging backings are exchange-owned and never enter the parcel ledger. Retirement is an exchange-local epoch, distinct from destination RAW/WAR tracking.

Designed: video-encoder exchange (foreign read continues after enqueue).

7. Schemes and GraphIR

Shipped. Public type: Scheme. Internally Goldy holds GraphIR — nodes, ownership-derived edges, group provenance, wave / partition analysis, retention fingerprints.

Scheme::include copies a child's description into the parent as one group (snapshot: later mutation of the child does not affect the parent; the child stays independently submittable). Scheme::group is sugar: a temporary child on the same Context, then include. Group-level .after(prior) expands to node-pair precedences when the schedule cache is rebuilt — never on the clean submit path.

Restrictions (all GoldyError::Validation at include time; the parent is left untouched): same Context; no pending record errors; no CPU dispatch, deposit, present/swapchain, yielding, or transient nodes; all child stamps alive (StaleResource otherwise).

On Scheme::submit:

  • Dependency analysis inserts barriers
  • Transient regions are colored for aliasing
  • Partitions may acquire exchange backing
  • Retained command buffers replay when bindings are unchanged

Designed scheme transformations (spec §8): fusion (mega-kernel), splitting at yield points (wavefront), dead-dispatch elision beyond basic analysis.

Goldy refuses ill-formed schemes (conflicting unordered private claims) rather than producing unspecified results — a deliberate narrowing of spec latitude.

8. Gates and ordering

Shipped. A gate is the interval between Scheme::submit calls and retirement via Context::boundary_crossed(T).

At a gate Goldy may reclaim deferred allocations, flush VRAM deferred rings, and service timeline signals.

Shipped cross-submission ordering: the runtime enforces ledger precedences across schemes on the same Context, using GPU barriers or host waits as needed. Clients must not assume which lever is used.

Pipeline depth (in-flight submissions) is client pacing — surface depth, FrameOrchestrator, when to consume claims. See Pipelined Frames.

9. Scripts: non-yielding today

Shipped. Public shaders today are non-yielding: [goldy_compute], [goldy_vertex], [goldy_fragment].

Designed yielding scripts:

  • $yield intrinsic in the virtual-entry-point transform
  • Script-state preservation (register spin-wait, workgroup-local, or parcel-backed)
  • Yield-point petitions (limited runtime power, not full gate powers)

10. Calling conventions

Shipped:

  • Slang as sole script language — Slang in One Source
  • Virtual entry points — typed parameters; virtual_main generates platform wrappers
  • Push-constant layout — bindless indices + scalars prepended per dispatch (backend-internal)
  • Access categories — validated at scheme record time

Designed: $yield petition descriptors, promise / continuation bindless category, paged-parcel fault servicing.

Portable programs depend on typed access categories and scheme structure, not on bindless heap layout.

11. Where Goldy constrains the spec

Deliberate restrictions for modern desktop / laptop workloads:

  • Refuses ill-formed schemes
  • Fixed Slang scripts
  • Closed typed-access category set
  • Workgroup-grid execution model
  • Single accelerator queue per device (heterogeneous multi-queue: Designed)
  • 2020+ hardware floor (Vulkan 1.4+, DX12 Enhanced Barriers, Metal Tier 2+)
  • Raster grain: one scheme node per framebuffer epoch (per-draw dispatches: Designed)

For older hardware or maximum portability, use wgpu. See Goldy vs wgpu and Target Hardware.

12. Abstract the medium, expose cost

Normative for Goldy design.

  • Layer A (medium): VRAM, residency, relocation, descriptor slots — abstracted; runtime-owned
  • Layer B (cost): Registers, occupancy, coalescing, access patterns, first-touch latency — exposed and queryable

Goldy must not present Layer A operations as uniform-cost or hide them entirely. Access-pattern types (Scattered vs Broadcast vs Interpolated) exist because hardware treats them differently.

Capability queries report backend, residency model, resize cost, zero-copy readback, and optional features honestly.

Appendix: Fondaco ↔ Goldy

Fondaco termGoldy / GPU analogue
SchemeScheme, GraphIR
DispatchCompute node; render-pass node (fused draws); copy / present
ScriptSlang shader (per pipeline); pass body is fused command list, not a Fondaco script
ParcelBuffer / Texture handle
MerchantProgram
ExchangeSurfaceExchange, MemoryExchange
Claim (exchange)Claim; deposit claims are Runtime-internal
Host claimPendingHostRead, HostView
GateFence epoch, boundary_crossed
WarehouseBudgetPolicy, VramAllocator
LedgerCross-submit sync analysis

Analogues are not equivalences. A scheme is the program's computation, not merely a scheduling artifact. An exchange preserves program sovereignty over parcels; a raw swapchain handle does not.

Full vocabulary: Terminology.

Render Passes and Schemes

Status: Research note. Complements Machine Specification and Goldy Runtime Mapping. Day-to-day recording: Render Pass Nodes.

Claim. In the Fondaco machine, a draw can be a dispatch. Goldy records many draws as one scheme node because Vulkan / DX12 / Metal in 2026 only expose a gate at the framebuffer epoch (the render pass), not at each draw. That clustering is a runtime transformation (spec §8 fusion), not a statement that draws are invisible to the machine.

Two grains

A dispatch is the finest unit the ledger can wait on: claims, gates, parcel retirement. A merchant may name a kernel launch, a copy, or a single raster draw as a dispatch.

A framebuffer epoch (native “render pass”, Goldy render_pass node) is the interval during which one or more attachments are checked out of the warehouse medium into on-chip tile / ROP storage. Load ops happen at checkout; store (and MSAA resolve, if any) happen at check-in. Other dispatches may not observe the attachment’s medium until check-in.

These are different objects:

DispatchFramebuffer epoch
MachineAtomic work; gate on either sideNot a machine primitive
2026 APIsKernel, copy, or a whole passvkCmdBeginRendering … EndRendering, Metal render encoder, DX12 BeginRenderPass
Waitable from outsideYes, if the runtime admits itYes — this is what the APIs actually fence
Private storageScript registers, groupsharedLive color/depth tile, bound pipeline / VB / IB

Goldy’s finish() is not a GPU opcode. It commits a fused dispatch: the epoch plus the draw list recorded inside it. Compute already terminates a node with dispatch(x, y, z) because that node is one launch. A raster node has no single “the” draw, so the terminator is explicit (Rust) or RAII (FFI / Python).

A pass has no native handle. Identity, if any, is the retained scheme node (or the command list it was recorded into). Vulkan 1.0 VkRenderPass objects had identity; Goldy sheds them (What Goldy Sheds). Dynamic rendering is two marks in a command stream.

Machine view: draws are dispatches

Spec §3: a dispatch runs to completion and is not reordered internally. Spec §8: the runtime may fuse adjacent dispatches if parcel states are preserved, and merchants should express natural grain.

Nothing in that forbids:

  1. Draw A (reads warm, writes attachment rt)
  2. Draw B (reads cool, writes rt)
  3. Compute C (overwrites warm)

with precedences A → B (both private on rt) and A → C (C private on warm). After A’s vertex/mesh stage retires, warm could be recycled while B still shades — if a gate existed after A that did not check rt back in.

The command processor (CP / firmware) already has that state machine: pipeline and VB binds are register writes; timestamp / event packets mark stage retirement per draw; TBDR hardware splits a pass into a binning job (vertex fetch for all draws) then per-tile fragment. Metal’s updateFence(afterStages: .vertex) is the one portable-ish leak of that vertex-job boundary. The machine is allowed to model it. The 2026 portable ABI is not required to.

Scripts remaining opaque (spec §4) still holds for each draw’s shader. It does not imply the sequence of draws is one script. Treating the sequence as opaque is Goldy’s lowering, not Fondaco ontology.

A kernel’s registers are private by physics. Pass-internal sticky state (current pipeline, current VB) is private by API contract. Drivers and firmware fiddle with it constantly; applications may not place a fence, event, or compute dispatch between two draws inside a Vulkan render pass instance. DX12 without BeginRenderPass is looser (interleaved compute + barrier) and thereby often flushes the tile. Variation across APIs is the proof that the intra-pass wall is policy, not the Fondaco machine.

What is physics

Tile residency of the attachment is the honest constraint. During the epoch, rt is not a stable VRAM image: on tilers it lives in GMEM; on immediate-mode GPUs it is spread across ROP caches and compression metadata. A compute dispatch that samples or overwrites rt mid-epoch forces a store/load — the same cost as splitting the pass.

So:

  • Fusing draws that share an attachment epoch is often required for cost, even on a runtime that could name each draw.
  • Fusing buffer claims (vertex buffers, bindless reads) with that epoch is not required by physics. It is required by Vulkan/Metal encoder rules: you cannot barrier warm and dispatch compute without ending the encoder, which checks rt in.

Recycling warm the moment draw A’s vertex stage is done, while draw B continues, is a real hardware opportunity. Portable 2026 APIs do not give Goldy a gate that releases warm without ending the epoch on rt. Practical lifetime for vertex memory is therefore frame-in-flight rings, not per-draw retirement.

Goldy view (shipped)

Goldy admits one render-pass node per epoch:

  • Claims are declared on the node (with_parcel, with_buffer_dependency, TargetLoad → public / private / private-inaugural on the target).
  • Body is a list of render commands: set pipeline, set VB/IB, draw, mesh dispatch, clear depth. Sequential sticky state. Not graph edges.
  • Gates land at node boundaries: after the native end-rendering, then copies, compute, present.

That is spec §8 fusion applied at record time because the substrate cannot address a finer gate. The scheme cannot insert a node between two draws, cannot fence one draw, and cannot start compute on warm until the whole pass’s graphics in the barrier’s source stages have retired — typically the entire pass, because the barrier is recorded after end-rendering.

with_parcel on the pass is the seam: the fused body still names parcels (set_vertex_buffer(warm)), so the program must echo those names as node claims or the ledger is a lie.

Designed (not shipped): per-draw dispatches plus a fusion hint “keep this attachment tile-resident.” A firmware-level or fondaco-native backend could split buffer retirement from attachment check-in. Until then, Goldy will not pretend a draw() is a waitable epoch.

Consequences for scheme authors

  • One-shot set_pipeline + draw(0..3) + finish() is a fused dispatch that happens to contain one statement. That is fine.
  • Two draws that must depth-test against each other must share an epoch (or the second pass Loads depth and pays a round trip).
  • Compute that consumes the color target is a later scheme node, after the pass — never inside the builder.
  • Overwriting a vertex parcel used by draw 1 while draw 2 of the same pass still runs is not expressible in Goldy 0.2. Split the pass (Load the target) if you need that gate; expect a store/load.

Design Thesis

Why Goldy exists, and how it differs from a conventional GPU library. Machine semantics live in Machine Specification; what is shipped today is in Goldy Runtime Mapping. Vocabulary: Terminology.

Executive summary

Goldy is a GPU runtime for the Fondaco Machine — programs own parcels (data) and express computation as schemes (dispatches + ownership-derived precedences). It targets modern native APIs (Vulkan 1.4+, DX12, Metal Tier 2+) with no translation layers, a single shader language (Slang), and a scheme-first API that sheds descriptor sets, explicit barriers, and swapchain ceremony.

The Fondaco model on GPU

Traditional GPU programming exposes descriptor set layouts, image layout transitions, render pass objects, pipeline layouts, and raw swapchain images with semaphores.

Fondaco instead gives the program:

ConceptRole
ParcelStable identity for data; physical medium is runtime-managed
SchemeFirst-class computation graph; precedences from ownership
ExchangeMediated foreign I/O (present, readback) via linear claims
GateWhere the runtime may relocate, reclaim, or insert work

Goldy's public API (Scheme, Parcel, SurfaceExchange, MemoryExchange) implements this model.

Access patterns, not graphics categories

Goldy names resources for what the hardware does, not which API invented the term:

Goldy termHardware behavior
ScatteredAny-thread read/write
BufRORead-only scattered (stronger cache hints)
BroadcastWave-broadcast constant fetch
InterpolatedDedicated texture filtering silicon
DirectSpatial2D/3D indexed access, no filtering
FilterSampler configuration

Shaders declare these as typed Slang parameters on [goldy_*] entry points. The CPU side uses matching buffer / texture kinds and scheme NodeAccess. See Parcels.

What Goldy sheds

Because Goldy requires modern baseline hardware, it drops:

Legacy conceptGoldy approach
Render pass objectsDynamic rendering
Descriptor set layoutsBindless (backend-internal)
Separate transfer queuesUnified queue model
OpenGL fixed functionShaders only
Multiple shader languagesSlang only

Details: What Goldy Sheds and Goldy vs wgpu.

Slang and virtual entry points

Goldy uses Slang as its sole shader language, compiled at runtime to SPIR-V / DXIL / MSL. The compiler is embedded in the crate.

import goldy_exp;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(MyUniforms cfg, Scattered<uint> data, ThreadId id) {
    data[id.x] = data[id.x] + cfg.base;
}

The virtual_main transform generates platform entry points with bindless slot resolution — see Virtual Entry Points. The goldy_exp standard library provides access functions, math, color utilities, vertex formats, and workgroup collectives.

Unified graphics and compute

Goldy treats graphics and compute as one scheme:

  • Compute simulation → raster present in one retained scheme
  • Compute-to-surface: compute writes swapchain drawables directly (no RenderPipeline) — Compute to Surface
  • Cross-scheme ordering via context timeline / ledger
  • Raster draws are machine-legal dispatches; Goldy fuses them into one pass node because 2026 APIs only fence at the framebuffer epoch — Render Passes and Schemes

This matches how modern engines and CUDA-style workloads converge on the same memory-access patterns. Goldy provides the primitives; performance patterns remain the developer's responsibility.

Backends

PlatformBackendNotes
WindowsDX12 (default), VulkanPIX on DX12
LinuxVulkanWayland surfaces
macOSMetalNative, not MoltenVK

Auto-selection with GOLDY_BACKEND override. Capability queries reflect backend-specific features honestly — Backend Architecture.

Goldy vs wgpu

wgpuGoldy
IdentityWebGPU for RustFondaco GPU runtime
GovernanceW3C specIndependent
Legacy floorVulkan 1.0+, web LCDVulkan 1.4+, modern only
Binding modelWebGPU bind groupsTyped bindless + schemes
BrowserYesNo (native only)

Use wgpu for web and maximum compatibility. Use Goldy for scheme-first Fondaco semantics and the modern feature union. Full write-up: Goldy vs wgpu.

Inspirations

SourceContribution
Sebastian Aaltonen — "No Graphics API"Target modern hardware; drop legacy ceremony
Ralph Levien — piet-gpu-hal post-mortemAbstract meaning, expose cost
Wayland compositor modelComplete frames, explicit sync, mediated present
SlangOne shader language, multi-backend
wgpuInstance / device ergonomics (adapted to schemes)
TU Darmstadt HAL paperMinimal necessary feature analysis

See also Motivation.

Roadmap posture

Shipped in 0.2.x: schemes, exchanges, compute-to-surface, growable buffers, retained replay, language bindings, Rust examples.

Designed: yielding scripts, scheme fusion / splitting, defragmentation, compute algorithm libraries (scan, sort, BLAS-class).

Speculative: WASI goldy-host, CUDA backend exploration, pre-2020 traditional-binding backend.

Do not treat Designed or Speculative items as shipped. Status table: Goldy Runtime Mapping.

Further reading

  1. Machine Specification
  2. Goldy Runtime Mapping
  3. Render Passes and Schemes
  4. Terminology
  5. Motivation · What Goldy Sheds · Target Hardware

Slang Quick Reference

Goldy uses Slang as its sole shading language. This page covers what you need to write Goldy shaders — not a full Slang language reference.

Basics

Slang uses HLSL-style syntax. If you've written HLSL or GLSL, most of it will look familiar.

Scalar Types

float f = 1.0;
int   i = -5;
uint  u = 10;
bool  b = true;

Vector and Matrix Types

float2 v2 = float2(1.0, 2.0);
float3 v3 = float3(1.0, 2.0, 3.0);
float4 v4 = float4(1.0, 2.0, 3.0, 4.0);

// Swizzling
float2 xy  = v4.xy;
float3 rgb = v4.rgb;

// Matrices
float4x4 mvp;
float4 transformed = mul(mvp, float4(pos, 1.0));

Structs

struct Particle {
    float2 position;
    float2 velocity;
    float  age;
};

Functions

float square(float x) { return x * x; }

// Public functions are exported from modules
public float3 my_effect(float2 uv) { return float3(uv, 0.5); }

Modules

Slang has a real module system (not #include). Modules are separate compilation units:

// In mylib.slang
module mylib;
public float3 effect(float2 uv) { return float3(uv, 1.0); }

// In shader.slang
import mylib;
float3 c = effect(uv);

goldy_exp Resource Types

The goldy_exp module defines type aliases that map to native Slang buffer and texture types. When used as parameters in [goldy_*] entry points, the Goldy compiler automatically resolves slot indices to live resource handles.

Buffer Types

Type aliasUnderlying typeAccess patternUsage
Scattered<T>StorageBuffer<T> (RWStructuredBuffer<T>)Read/write, any thread, any addressdata[i], data[i].field = v
BufRO<T>ReadOnlyBuffer<T> (StructuredBuffer<T>)Read-only, hardware read-cache hintdata[i]
ByteAddressByteAddressView (RWByteAddressBuffer)Raw byte-level access.Load(addr), .Store(addr, v), .InterlockedMin(...)

Texture Types

Type aliasUnderlying typeAccess patternUsage
Interpolated<T>Texture2D<T>Hardware-filtered samplingtex.Sample(samp, uv), tex.Load(loc)
DirectSpatial<T>RWTexture2D<T>Direct 2D read/write, no filteringimg[int2(x,y)], img.GetDimensions(w,h)

Sampler Type

Type aliasUnderlying typeUsage
FilterSamplerStatePass to tex.Sample(filter, uv)

Broadcast (Constant Buffer)

To pass uniform data (same value for all threads), declare a struct type directly as a parameter — no wrapper needed. The codegen recognizes any non-resource, non-system-value struct as a constant-buffer broadcast:

struct TimeUniforms { float time; float delta_time; };

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(TimeUniforms cfg, Scattered<Particle> particles, ThreadId id) {
    particles[id.x].position += particles[id.x].velocity * cfg.delta_time;
}

System-Value Types

Declare these as parameters in [goldy_*] entry points to receive GPU-provided values. The codegen maps each type to its SV_* semantic automatically.

Compute

TypeMaps toComponents
ThreadIdSV_DispatchThreadID.x, .y, .z, .xy, .xyz
GroupThreadIdSV_GroupThreadID.x, .y, .z, .xy, .xyz
GroupIdSV_GroupID.x, .y, .z, .xy, .xyz

Graphics

TypeMaps toComponents
VertexIdSV_VertexID.value
InstanceIdSV_InstanceID.value
IsFrontFaceSV_IsFrontFace.value

Ray tracing

TypeMaps toComponents
DispatchRaysIndexSV_DispatchRaysIndex.x, .y, .z
DispatchRaysDimensionsSV_DispatchRaysDimensions.x, .y, .z

Use [goldy_raygen] / [goldy_miss] / [goldy_closesthit] with Accel and an inout payload struct on miss and closest-hit. Dispatch with Scheme::trace_rays (ray counts, not workgroups).

Entry Point Attributes

[goldy_compute]

Marks a compute shader entry point. The Goldy compiler generates the real [shader("compute")] wrapper that resolves resource slots and system values.

import goldy_exp;

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<uint> data, uint offset, ThreadId id) {
    data[id.x + offset] += 1;
}

[goldy_vertex]

Marks a vertex shader entry point.

import goldy_exp;

struct VSOutput {
    float4 position : SV_Position;
    float4 color    : COLOR;
};

[goldy_vertex]
VSOutput vs_main(BufRO<Vertex> verts, VertexId vid) {
    Vertex v = verts[vid.value];
    VSOutput o;
    o.position = float4(v.pos, 0.0, 1.0);
    o.color    = v.color;
    return o;
}

[goldy_fragment]

Marks a fragment shader entry point.

import goldy_exp;

[goldy_fragment]
float4 fs_main(Interpolated<float4> tex, Filter samp, float2 uv : TEXCOORD0) : SV_Target {
    return tex.Sample(samp, uv);
}

[goldy_raygen] / [goldy_miss] / [goldy_closesthit]

Ray-tracing pipeline stages. Miss and closest-hit take an inout payload; raygen uses bindless resources like compute. The shader-binding table is owned by RayTracingPipeline.

import goldy_exp;
struct HitPayload { uint hit; }

[goldy_raygen]
void rgen_main(Accel scene, Scattered<uint> hits, DispatchRaysIndex idx) {
    RayDesc ray;
    ray.Origin = float3(0, 0, -2);
    ray.TMin = 0.001;
    ray.Direction = float3(0, 0, 1);
    ray.TMax = 100;
    HitPayload p;
    TraceRay(scene, RAY_FLAG_FORCE_OPAQUE, 0xFF, 0, 0, 0, ray, p);
    hits[idx.x] = p.hit;
}

[goldy_miss]
void rmiss_main(inout HitPayload p) { p.hit = 0; }

[goldy_closesthit]
void rchit_main(inout HitPayload p) { p.hit = 1; }

Common Patterns

Accessing Buffers by Index

All Scattered<T> and BufRO<T> parameters support standard array indexing. Field-level writes work directly on Scattered<T>:

[goldy_compute]
[numthreads(64, 1, 1)]
void cs_main(Scattered<Particle> particles, ThreadId id) {
    Particle p = particles[id.x];
    p.position += p.velocity;
    particles[id.x] = p;

    // Or field-level write:
    particles[id.x].age += 1.0;
}

Sampling Textures

[goldy_fragment]
float4 fs_main(Interpolated<float4> albedo, Filter samp, float2 uv : TEXCOORD0) : SV_Target {
    return albedo.Sample(samp, uv);
}

Writing to Storage Images

[goldy_compute]
[numthreads(8, 8, 1)]
void cs_main(DirectSpatial<float4> output, ThreadId id) {
    output[int2(id.x, id.y)] = float4(float(id.x) / 512.0, float(id.y) / 512.0, 0.5, 1.0);
}

Fullscreen Triangle (Vertex-less)

Use vs_fullscreen_triangle() from goldy_exp to render fullscreen effects without a vertex buffer:

import goldy_exp;

[shader("vertex")]
FullscreenVarying vs_main(uint vertex_id : SV_VertexID) {
    return vs_fullscreen_triangle(vertex_id);
}

[shader("fragment")]
float4 fs_main(FullscreenVarying input) : SV_Target {
    return float4(input.uv, 0.5, 1.0);
}

Compute + Render Buffer Sharing

Compute shaders and graphics shaders share the same bindless buffers. The scheme handles the dependency:

// Compute: update particles
[goldy_compute]
[numthreads(64, 1, 1)]
void cs_update(TimeUniforms cfg, Scattered<Particle> particles, ThreadId id) {
    particles[id.x].position += particles[id.x].velocity * cfg.delta_time;
}

// Vertex: read particles for rendering
[goldy_vertex]
VSOutput vs_draw(BufRO<Particle> particles, InstanceId iid, VertexId vid) {
    Particle p = particles[iid.value];
    // Generate quad geometry from particle position...
}

Rust-Side Resource Binding

Resources are bound in declaration order (left to right in the shader signature) via with_parcel. Graph access (NodeAccess) drives dependency analysis; the descriptor kind (SRV vs UAV) is chosen from pipeline reflection at dispatch / set_pipeline:

#![allow(unused)]
fn main() {
scheme
    .node("update", &pipeline)
    .with_parcel(&cfg_buf, NodeAccess::Read)
    .with_parcel(&particle_buf, NodeAccess::ReadWrite)
    .dispatch(workgroups, 1, 1);
}

Use NodeAccess::Overwrite when a compute node fully replaces a parcel without reading prior contents (ping-pong write sides, clears). Plain scalar parameters (uint offset) are also push-constant bindings — no wrapper struct needed.

goldy_exp Utility Modules

ModuleContents
goldy_exp/math.slangPI, TAU, hash(), hash2(), center_uv(), scale_uv(), to_polar(), smootherstep()
goldy_exp/color.slangrainbow(), palette(), heat(), hsv_to_rgb(), luminance(), gamma_correct()
goldy_exp/primitives.slangquad_position(), quad_position_rotated(), billboard_position(), fullscreen_position(), fullscreen_uv()
goldy_exp/types.slangParticle2D, Particle3D, FrameUniforms, Transform2D, DispatchShape
goldy_exp/vertex.slangFullscreenVarying, ColoredVertex, ColoredVarying, vs_fullscreen_triangle()
goldy_exp/access.slangResource type aliases and system-value types (documented above)
goldy_exp/interlocked.slangInterlocked<T> cells; InterlockedLoad/Store/Add/Or/Xor/Min/Max/Exchange

Further Reading

Environment Variables

Goldy reads several environment variables at runtime for backend selection, validation, debugging, and Slang configuration.

General

VariableValuesDefaultDescription
GOLDY_BACKENDvulkan, vk, dx12, d3d12, directx, metal, mtl, cuda, webgpu, wgpu, cpuPlatform default (macOS → Metal, Windows → DX12, Linux → Vulkan)Override backend selection at runtime. Shipped: Vulkan, DX12, Metal. In progress: CUDA, WebGPU. cpu is a compute-only host-callable JIT path (never a platform default).
GOLDY_MATMULnative, fallback (stdlib, goldy)nativeSelect the MatMul implementation. native uses cuBLAS on CUDA and MPS on Metal, and falls through to the Goldy stdlib kernel on other backends. fallback always uses the portable stdlib kernel (for differential tests).
GOLDY_SLANG_PATHFile path(not set)Override the path to the Slang shared library (slang.dll / libslang.dylib / libslang.so). Bypasses the default search order (vendored next to executable → extracted from embedded).
GOLDY_FFI_PATHFile path(not set)Full path to the goldy_ffi shared library (goldy_ffi.dll / libgoldy_ffi.so / libgoldy_ffi.dylib). Used by goldy-ffi-client for runtime library loading.

Validation

VariableValuesDefaultDescription
GOLDY_VALIDATIONComma/semicolon/whitespace-separated list: api, layout, layouts, host_access, scheme, graph, readback, timeline, all; or 1 / true / yes(not set)Enable validation categories. api enables GPU API validation (Vulkan validation layers + debug messenger, Metal shader validation, CUDA Driver diagnostics: PTX JIT logs, eager stream sync, launch-limit checks; WebGPU/wgpu validation error scopes on shader/PSO create and bind groups). layout enables Rust/Slang struct layout and buffer stride checks. host_access page-protects CPU-visible GPU copies (CPU backend parcels; slower, not complete). scheme / graph / readback enable retained host-read staging checks and strict graph lifetime checks (Accel must be built in the same scheme before RayQuery / TraceRays). Cycle detection, mesh vs draw mix-ups, BLAS/TLAS misuse, and missing BufferFlags::ACCEL_INPUT always run on submit with GoldyError::Validation messages that include a hint:. all enables layout, api, timeline, scheme, and host_access (static shader checks are separate: GOLDY_SHADER_VALIDATION). The shorthand 1 / true / yes enables GPU API only (layout stays opt-in). Deep CUDA memory/race checking is not covered — use external compute-sanitizer.
GOLDY_VALIDATION_FATAL1, true, yes(not set)Separate from GOLDY_VALIDATION. When GPU API validation is on, treat Vulkan Khronos ERROR messages as hard failures (Err on later Goldy Result calls; panic on backend drop so cargo test fails). Without this, messages are logged (goldy::validation) and successful vk* calls still succeed. WebGPU validation scopes already return Err without this flag.
GOLDY_VALIDATE_LAYOUTS1, true, yes(not set)Legacy toggle for layout validation only. Equivalent to GOLDY_VALIDATION=layout.
GOLDY_SHADER_VALIDATIONList of check names processed left to right: all, bounds, -bounds (exclude), none; or 1 / true / yes for all(not set)Static checks over Slang's front-end IR at shader compile time. Independent of GOLDY_VALIDATION and not implied by GOLDY_VALIDATION=all: each shader is compiled a second time to IR and analyzed whole-program, and findings mean "not proven" rather than "wrong". bounds is the static bounds analysis: every dynamic array / vector / matrix index it cannot prove inside 0 <= index < length is logged as a warn with its Slang source location and call path. Never fails a compile.

Validation Examples

# GPU API validation (Vulkan layers, Metal shader validation, CUDA Driver diagnostics)
GOLDY_VALIDATION=api cargo run --example triangle

# Layout + stride checks only
GOLDY_VALIDATION=layout cargo run --example triangle

# Everything (runtime validation)
GOLDY_VALIDATION=all cargo run --example triangle

# Static shader checks are a separate switch (warn on unproven dynamic array indices)
GOLDY_SHADER_VALIDATION=bounds cargo run --example triangle

# Fail Goldy calls / tests when Vulkan records an ERROR
GOLDY_VALIDATION_FATAL=1 GOLDY_VALIDATION=all cargo test --features vulkan
GOLDY_VALIDATION_FATAL=1 GOLDY_VALIDATION=api cargo test --features vulkan

# Shorthand for GPU API only
GOLDY_VALIDATION=1 cargo run --example triangle

# CUDA with API validation (JIT logs, eager sync, launch limits)
GOLDY_BACKEND=cuda GOLDY_VALIDATION=api cargo test --features cuda

DX12-Specific

VariableValuesDefaultDescription
GOLDY_DX12_DEBUG1, trueOn in debug buildsEnable the D3D12 debug layer. On by default in debug builds; set explicitly for release builds.
GOLDY_DX12_NO_DEBUG1, true(not set)Force-disable the D3D12 debug layer even in debug builds. Useful to avoid debug-layer crashes in parallel test threads.
GOLDY_DX12_GBV1, true(not set)Enable D3D12 GPU-Based Validation. Catches UAV/SRV descriptor mismatches, resource state errors, and out-of-bounds access on the GPU timeline. Very slow — use for targeted debugging only.
GOLDY_DX12_FORCE_WARP1, true(not set)Force the DX12 backend to use the WARP software rasterizer, even when hardware GPUs are present. Use for headless CI or reproducing WARP-specific rendering bugs.
GOLDY_DX12_ALLOW_WARP1, true(not set)Allow the WARP adapter to appear in device enumeration. Without this or GOLDY_DX12_FORCE_WARP, WARP is hidden.

Debugging

VariableValuesDefaultDescription
GOLDY_DUMP_SHADERSDirectory path(not set)Dump compiled shaders to the specified directory at compile time. Vulkan: {entry}_h{handle}_vulkan.spv. DX12: {entry}_h{handle}_dx12.dxil. Metal: {idx}_{entry}.metal. CUDA: {entry}_h{handle}_{spec}_cuda.cu (Slang CUDA C++) and .ptx (NVRTC/Slang PTX loaded by the CUDA driver), plus goldy_apply_dispatch_shape.{cu,ptx} for the graph updater.
GOLDY_DUMP_RUST_KERNELS1 / true / directory path(not set)Dump canonical [goldy_compute] Slang and structured ABI metadata produced by #[goldy::compute] during Kernel::prepare. 1/true writes under the process temp dir (goldy_rust_kernels/).
GOLDY_CPU_SHADERS1, true, yes(not set)Documented gate for the standalone goldy::cpu_shaders APIs. GPU backends ignore this. Scheme submit uses GOLDY_BACKEND=cpu instead.
GOLDY_GPU_PROFILEAny non-empty value; optional chrome[=path](not set)Enable GPU timestamp profiling logs. On Vulkan/DX12, records per-dispatch GPU durations. On Metal, records command-buffer GPU duration. chrome / chrome=/path.json also writes a Perfetto Chrome-trace JSON file. Disables retained CB reuse while active.
GOLDY_DISABLE_CB_REUSE1, true, yes(not set)Disable retained command-list reuse for schemes: every submit re-records through ordinary graph submission. Useful to isolate a bug to the retention path.
GOLDY_SPECIALIZATION0, false, no, off to disableOnRetained-scheme shader specialization prediction. When on, a compute dispatch whose with_param scalars hold still across clean submits is moved to a variant of its shader with those words baked in (a full recompile on a worker thread, swapped in as a params-only re-record, demoted the moment a baked word changes). Output is identical either way; ReplayStats::specialization_* and Scheme::node_is_specialized show what happened. Pays off when the baked word gates real work, not when it only feeds arithmetic. Has no effect on WebGPU or the CPU backend, which decline via a backend capability. See What baking actually compiles.
GOLDY_SHADER_TIMING1 / any value other than 0(not set)Print stderr wall-clock breakdown of Slang cache lookup, search-path hashing, and PSO create during shader compile
GOLDY_METAL_CAPTURE1 / path / path,skip=N,frames=M(not set)Metal only. Opt-in programmatic MTLCaptureManager GPU capture for Xcode Metal Debugger. 1/true/yes captures to Developer Tools; a path writes a .gputrace. Optional skip=N (default 60) skips warm-up submits; frames=M (default 1) captures M submits. Automatically sets METAL_CAPTURE_ENABLED=1 if unset. Open the .gputrace in Xcode → Performance to inspect register pressure, occupancy, and per-line shader costs.
GOLDY_API_LOGFile path(not set)Metal only. Append NDJSON Metal API call traces (dispatches, encoder open/close, commits) to the given file.
GOLDY_API_LOG_SYNC1(not set)Force synchronous GOLDY_API_LOG writes (for tests).

Metal capture example

# Warm up 120 submits, then write one .gputrace for Xcode Metal Debugger
GOLDY_METAL_CAPTURE=/tmp/capture-tiger.gputrace,skip=120,frames=1 \
  target/release/with_winit_bin --timeout-secs 12 --no-vsync

# Open /tmp/capture-tiger.gputrace in Xcode → click Performance → select fine_area

Interop with System Variables

Goldy also respects these non-Goldy environment variables:

VariableBackendDescription
VK_INSTANCE_LAYERSVulkanIf set to include VK_LAYER_KHRONOS_validation, Goldy enables Vulkan validation regardless of GOLDY_VALIDATION.
VK_LAYER_PATHVulkanStandard Vulkan loader variable for locating validation layer manifests.
MTL_SHADER_VALIDATIONMetalWhen GOLDY_VALIDATION enables API validation and this variable is unset, Goldy sets it to 1 before creating the first Metal device. If you set it yourself, Goldy does not override it.
METAL_CAPTURE_ENABLEDMetalRequired for programmatic GPU capture outside Xcode. Goldy sets this to 1 automatically when GOLDY_METAL_CAPTURE is set (if unset).
CUDA_LAUNCH_BLOCKINGCUDAWhen GOLDY_VALIDATION enables API validation and this variable is unset, Goldy sets it to 1 before CUDA driver init. If you set it yourself, Goldy does not override it. Forces synchronous kernel launches so errors surface at the launch site.
WGPU_BACKENDWebGPUwgpu instance backend mask (vulkan, metal, dx12, gl, …). Goldy passes this through wgpu::InstanceDescriptor::from_env_or_default(). CI uses vulkan on Linux, metal on macOS, and dx12 on Windows.
WGPU_FORCE_FALLBACK_ADAPTERWebGPUWhen set (1/true), wgpu prefers a software adapter (WARP on Windows). Used in Windows CI because hosted runners have no discrete GPU.

License

Goldy is licensed under the MIT License.

MIT License

You may use Goldy freely in any project — including proprietary and commercial software:

  • Use, modify, and distribute Goldy
  • Static or dynamic linking in proprietary software
  • No obligation to release your own source code
  • Include the MIT copyright notice and license text in distributions

See LICENSE in the repository for the full text.

Dependencies

Goldy depends on various open-source libraries with their own licenses:

DependencyLicense
ashMIT/Apache-2.0
anyhowMIT/Apache-2.0
thiserrorMIT/Apache-2.0
tracingMIT
bitflagsMIT/Apache-2.0
bytemuckZlib/MIT/Apache-2.0

All dependencies are permissively licensed.