AI Engineer, Recursive Self-Improvement for Compute
DevExplore · Santa Clara, CA · 6 days ago
EngineeringFull-time
About the role
The Role & The Person We are hiring AI Engineers to build recursive self-improvement systems for compute. This role sits at the intersection of AI systems, performance engineering, hardware-aware optimization, and agentic software development.
Key Responsibilities
- Build agentic and learning-driven optimization loops for compute workloads and hardware engineering workflows
- Develop automated systems that generate, compile, test, benchmark, profile, and iterate on candidate improvements with minimal human intervention
- Collaborate with AI researchers on reward design, reward shaping, reward hacking analysis, long-horizon optimization, and model improvement loops
- Design feedback systems that accumulate useful data from successful attempts, failed attempts, profiler traces, benchmark results, and validation logs
- Improve iteration speed through staged validation, caching, parallel execution, proxy metrics, and faster feedback paths
- Build reusable tools and patterns that can generalize across multiple compute and hardware optimization domains
- Mentor other engineers and help set technical direction for self-improving AI systems for compute
Technical Focus Areas
- AI-assisted program optimization, code generation, debugging, and automated repair
- GPU and CPU performance engineering, including kernels, libraries, compilers, profilers, and benchmark-driven optimization
- Agentic workflows that use tools, tests, simulators, profilers, and structured feedback to improve over time
- Reinforcement learning, post-training, reward modeling, and evaluation methods for engineering tasks
- Hardware-aware optimization where correctness, latency, throughput, area, power, timing, or resource usage may all matter
- Experiment platforms, leaderboards, dashboards, and data pipelines for reproducible comparison and continuous improvement
Qualifications & Experience
- Education: Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or a related field (or equivalent practical experience)
- Master's preferred; PhD is a plus, especially in AI systems, reinforcement learning, compilers, GPU computing, or hardware/software co-design
- Required Skills: Strong software engineering experience in Python and at least one systems language such as C++, C, HIP, or CUDA
- Experience building AI, ML, agentic, optimization, or automation systems that are evaluated with objective metrics
- Ability to design reliable experiment loops, benchmark harnesses, validation workflows, and correctness/performance evaluation pipelines
- Strong technical judgment in debugging, profiling, root-cause analysis, and performance-oriented iteration
- Clear communication skills and ability to work across AI research, hardware, software, and partner-facing teams
Preferred Experience
- Experience with GPU kernels, ROCm/HIP, CUDA, Triton, PyTorch, JAX, TensorFlow, or distributed training/inference systems
- Experience with reinforcement learning, post-training, reward modeling, automated program optimization, or agentic coding systems
- Familiarity with CPU performance engineering, compiler optimization, benchmarking, profiling, or math libraries
- Exposure to hardware design, simulation, formal verification, performance/power/area analysis, or hardware/software co-design
- Experience building production-quality evaluation platforms, experiment tracking, dashboards, or leaderboards
- Publications, open-source contributions, or shipped systems in AI systems, GPU computing, compilers, RL, or hardware/software co-design are a plus