C3D is a small graphics API for 2D/3D-style rendering with selectable NVIDIA CUDA and OpenCL backends. It was built for fun and experimentation, not as a serious production renderer or a replacement for a native graphics API.
It exposes a compact graphics-like interface for:
- stage buffers
- textures
- vertex and index buffers
- command buffers
- render passes
- line, quad, and triangle draws
- swapchain
The repo also includes a Windows demo in test/main.cpp that renders into a backend texture, copies the result back to the CPU, and presents it with GDI.
Several write and render-path calls also support internal object cycling through the cycle flags. The behavior is intended to work like SDL_gpu: the API can rotate to a fresh internal backing resource to avoid stomping data that may still be in flight on the GPU.
I built C3D to experiment with a graphics API shape without depending on a full native graphics stack like Direct3D, Vulkan, or OpenGL.
The idea is to keep the public API small and C-friendly, but run the heavy lifting on a GPU backend that can be switched at build time.
This project is intentionally lightweight and exploratory. The goal is to learn, prototype, and test ideas, not to chase production readiness.
inc/C3D.h: public APIsrc_cuda/*.cu: CUDA implementationsrc_opencl/*.cpp: OpenCL implementationsrc_opencl/render_kernels.cl: OpenCL rasterizer kernels compiled at runtimetest/main.cpp: Windows demo applicationproject.bbs: build description for thebbsbuild system
The public header is organized around a few resource types plus a minimal command submission model:
graph TD
Stage["Stage Buffers"] <--> Textures["Textures"]
Textures <--> Swap["Swapchain"]
Buffers["GPU Buffers"] <--> Cmd["Command Buffers"]
Cmd <--> Textures
Stage <--> Buffers
Pass["RenderPass"] --> Cmd
To build the project, install bbs.
Then pick a backend config and run bbs build from the repository root.
The OpenCL headers and ICD loader are fetched automatically from Khronos through the opencl bbs dependency. You do not need to install an OpenCL SDK. You do need an OpenCL runtime supplied by your graphics driver; NVIDIA's driver/CUDA installation already provides one on this machine.
CUDA builds:
bbs build -t c3d_test -c cuda-debug
or
bbs build -t c3d_test -c cuda-release
OpenCL builds:
bbs build -t c3d_test -c opencl-debug
or
bbs build -t c3d_test -c opencl-release
Profiling builds with Tracy:
bbs build -t c3d_test -c cuda-debug-profile
bbs build -t c3d_test -c cuda-release-profile
bbs build -t c3d_test -c opencl-debug-profile
bbs build -t c3d_test -c opencl-release-profile
The same c3d_lib and c3d_test targets are reused for every configuration. CUDA configs compile src_cuda, OpenCL configs compile src_opencl, and the public inc/C3D.h API stays unchanged across both backends.
Right now the main bottleneck is Windows presentation through GDI.
C3D renders into CUDA-managed GPU memory, but the demo cannot present that memory directly to a window surface. Instead, it has to copy the rendered image back to CPU-visible memory and hand it off to GDI for presentation. That extra GPU-to-CPU transfer plus the GDI blit is currently the slowest part of the frame path.
Here is a Tracy capture from the demo showing the presentation path under inspection:

