The Opportunity
Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure.
Responsibilities
- Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads.
- Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks.
- Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals.
- Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries.
- Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity.
- Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines.
- Benchmark against CPU/GPU baselines and drive continuous performance improvements.
- Contribute to internal libraries, documentation, and best practices for GPU performance engineering.
- 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience.
- Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy).
- Proficiency in C++ and CUDA; experience with Python for tooling/orchestration.
- Experience with performance profiling tools (Nsight Systems/Compute).
- Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand).
- Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design.
- Comfortable working across the stack from low-level kernels to system-level infrastructure.
Nice to Have
- Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks.
- Exposure to data center power/thermal management or demand-response systems.
- Background in HPC, quantitative finance, or large-scale distributed systems.
- Familiarity with Kubernetes/Slurm for GPU cluster orchestration.
- Interest or experience in energy markets, grid systems, or sustainability-focused compute.
- Competitive salary and an equity sign-on bonus.
- Biannual bonus scheme.
- Fully expensed tech to match your needs.
- Breakfast and dinner allowance for office based employees.