C++ CUDA Engineers | $100/hr
The Ai Training Company · United States · Today
RemoteRemoteEngineeringFull-time
GPU Kernel Optimization Engineers. Remote AI Project You will analyze, profile, and optimize GPU kernels across modern hardware environments using C++17, Python, CUDA, HIP, shaders, and related GPU programming technologies. What You’ll Do Analyze and optimize GPU kernels for speed, efficiency, memory usage, and hardware utilizationIdentify performance bottlenecks using profiler-guided analysisUse metrics such as occupancy, L2 cache hit rate, L2 throughput, memory bandwidth, latency, and compute utilizationReview and improve existing C++, Python, CUDA, HIP, shader, or kernel implementationsOptimize kernels without requiring deep prior knowledge of every underlying algorithmClearly document optimization decisions and performance tradeoffs Who Can Apply Relevant backgrounds include: GPU Kernel Engineers, GPU Software Engineers, CUDA Developers, CUDA Engineers, GPU Performance Engineers, GPU Computing Engineers, HPC Engineers, High-Performance Computing Developers, Parallel Computing Engineers, Graphics Engineers, Rendering Engineers, Shader Engineers, GPGPU Engineers, Compiler Engineers, Performance Optimization Engineers, Systems Performance Engineers, Machine Learning Systems Engineers, AI Infrastructure Engineers, Deep Learning Performance Engineers, ML Compiler Engineers, Inference Optimization Engineers, Training Optimization Engineers, Embedded GPU Engineers, Scientific Computing Engineers, and Research Engineers. Professionals with experience in CUDA, HIP, ROCm, Slang, HLSL, GLSL, OpenCL, SYCL, Vulkan Compute, Metal Shading Language, Triton, PTX, tensor cores, GPU compilers, parallel programming, or heterogeneous computing are encouraged to apply. RequirementsStrong command of C++ through C++17Working knowledge of Python and GitFluency in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, OpenCL, SYCL, or a related technologyAt least 1 year of professional or graduate-level GPU programming or research experienceExperience profiling and optimizing GPU kernelsUnderstanding of GPU architecture, memory hierarchy, occupancy, parallelism, and performance metrics Preferred ExperienceCUDA C++ Core LibrariesInline PTX assemblyTensor core-level optimizationNVIDIA Blackwell, Hopper, or Ampere architecturesAMD GPU or ROCm optimizationNVIDIA Nsight Compute or similar GPU profiling toolsTriton kernel developmentGPU compiler or ML compiler optimizationOpen-source GPU kernel contributionsPrevious experience with NVIDIA, AMD, Qualcomm, Intel, or another GPU hardware organization This opportunity is ideal for specialists who enjoy profiling low-level code, identifying hardware bottlenecks, and extracting maximum performance from modern GPU architectures.