478. The GPU API is confined to one thread, scoped by pass, and checked in Java
Date: 2026-09-24
Status
Accepted. The public half of docs/gpu-plan.md’s phase 2, over the SDL_GPU
wrappers of ADR-0475 and ADR-0476.
Context
Phase 2 asks for “a safe layer an application can use without knowing
SDL_GPU’s structs”. canvas3d renderers are its first users (phase 5), then
the composited window’s layers and video (phases 3, 4 and 6). They need
textures, buffers, samplers, shaders, pipelines, a command scope per frame,
uploads, and readback.
The wrappers in natives.sdl.gpu already check SDL’s rules in Java: one pass
at a time, a draw with everything bound, resources of the same device. But they
are :natives types, and ADR-0280 keeps those out of every signature an
application can read. They also cover only what the toolkit’s quads needed.
Vertices came from the vertex id, there was no depth, and a pipeline was always
a quad pipeline.
SDL_GPU has sharp edges a Java API can blunt:
- A command buffer cannot be cancelled with a pass open.
- A debug group begun in a pass has to end in that pass on Metal.
- A texture cycled on upload loses whatever the upload did not write.
- A transfer buffer mapped for writing makes the CPU wait if the GPU still reads it, unless it is cycled.
- Nothing is thread-safe.
Decision
A public package, io.github.digitalsmile.goldberry.gpu, whose types wrap
the natives wrappers. No SDL type, handle or MemorySegment appears in it.
- Resources.
GpuTexture,GpuBuffer,GpuSampler,ShaderandGraphicsPipelineextend a sealedGpuResource. Each isAutoCloseableand made from a record:TextureSpec,SamplerSpec,ShaderCode, orPipelineSpecwith its builder. The enums those records name map onto SDL’s by exhaustiveswitch, and a test holds each mapping to be one to one. A use after close, or with another device, throws in Java and names the public type. - One thread. A
GpuDevicebelongs to the thread it was handed out on. Every call from another throwsWrongThreadException, as a confinedArenadoes, before the driver is reached. - The device is not an application’s to make or close.
GpuDevice.wrapis package-private. The backend will own the one device of the process (D2) and hand it out throughgpuSurface()(phase 3). - Passes are scoped by lambdas.
GpuFrame.copyPass(body)andrenderPass(target, load, [depth,] body)open the pass, run the body, and end the pass whether it returns or throws. Passes do not nest, and a pass kept past its body throws. AGpuFrameisAutoCloseable: closing an unsubmitted frame discards it, and closing a submitted one does nothing, so try-with-resources is always right.debugGroup(name, body)is scoped the same way, and the wrapper refuses a group left open across a pass’s end. - Uploads go through one staging buffer per device. It starts at 64 KiB and at least doubles when an upload does not fit. Each map is cycled, so writing never waits for the GPU and never overwrites bytes it has yet to copy. An upload of regions records the texture upload uncycled, so the pixels outside the regions stay as they were; that is the composited UI’s damage-only upload. An upload of the whole texture or buffer is cycled.
- Readback is recorded in the frame and awaited after it.
GpuFrame.readback(texture, region)records a download. Submitting a frame with readbacks takes one fence, which they share and release.Readback.await()is the only blocking call in the API.awaitPixels()gives aPixelBufferof premultiplied BGRA, swizzling an RGBA texture. - Driver refusals are
GpuException. What the API can see for itself isIllegalArgumentExceptionorIllegalStateException.
Under it, the natives wrappers grow what canvas3d needs, each piece checked
the same way:
- vertex and index buffers, and their uploads and downloads;
- indexed and instanced draws;
- a
SdlGpuPipelineDescriptionwith vertex input, primitive type, culling, front face and a depth test; - depth targets for render passes;
- sampler address modes;
- uniforms pushed as bytes;
- balanced debug groups.
That is ten more exports (SDL_CreateGPUBuffer … SDL_InsertGPUDebugLabel)
and seven more verified structs. The enumerators are verified through nine new
enums, so the names and values are checked against the compiled library.
Alternatives considered
- Exporting the natives wrappers. They are already safe, but it would
reverse ADR-0280 for one module. It would also make
SdlGpu…names, and every change to them, part of the published surface. - Passes as
AutoCloseableobjects returned to the caller. Safe inside try-with-resources, but nothing makes a caller use one, and a pass left open locks the command buffer. With a lambda there is nothing to forget. The cost is that state a body computes has to leave through a captured variable. - One transfer buffer per upload. Simpler, but a 4K frame would allocate 33 MB every frame. The staging buffer is sized once, and SDL’s cycling keeps the few copies frames in flight need.
awaitreturning the transfer buffer’s mapped memory. It would need no copy, but the memory would stop being valid at unmap, the lifetime ADR-0019 keeps raw memory from having. The copy goes to the heap, and only readback pays for it.
Consequences
canvas3d(phase 5) has what its cube needs: vertex and index buffers, depth, culling and a uniform block. The tests draw a mesh with shaders of their own, which:gpu:compileShadersnow compiles fromsrc/test/shadersinto test resources that do not ship.ShaderManifestTestholds them to their sources too.- Phase 2’s exit is met on Metal: a triangle from a vertex buffer reads back as a reference rasterized in Java draws it, every pixel but those on an edge. Lavapipe waits for the GPU lane.
- The composited window (phase 3) builds on
GpuFrame. Its swapchain becomes a secondRenderTarget(the interface is sealed and permits onlyGpuTexturetoday), andCopyPass.upload(texture, pixelBuffer, damage)is its UI upload.SDL_CancelGPUCommandBufferis refused after a swapchain acquire, so a composited frame will have to submit rather than discard. - Vertex-buffer and sampler bindings are reset whenever a pipeline is bound. SDL keeps them on some backends. The stricter rule is the same on all of them.
- Not in this cut, and recorded so it is not rediscovered:
PresentMode, which has no consumer before phase 3;- resource names (
SDL_SetGPUTextureName); - mip levels, multisampling and stencil;
- compute (phase 7).
- The culling test checks winding in normalised device coordinates, y up, as the API documents. If lavapipe disagrees with Metal, that is SDL’s backends disagreeing, and the test is where it will show.