DevOps Engineer II
About D‑Wave
D-Wave is a leader in the development and delivery of quantum computing systems, software, and services. It is the world’s first commercial supplier of quantum computers, and the first and only to offer dual‑platform quantum computing products and services, spanning both annealing and gate‑model quantum computing technologies. D-Wave's innovative solutions enable organizations across various industries to explore new computational possibilities, solve complex problems, and accelerate scientific discovery. With a commitment to advancing the field of quantum computing, D-Wave continues to push the boundaries of technology and innovation, fostering a collaborative environment that encourages growth and development.
About The Role
We are seeking a DevOps Engineer specializing in Platform Reliability & Observability to join our dynamic team. Reporting directly to the DevOps Engineering Manager, this role is pivotal in supporting and maintaining hybrid infrastructure platforms that include on‑premises environments, Kubernetes clusters, cloud services, and CI/CD pipelines. The ideal candidate will have a strong background in systems engineering, infrastructure automation, and cloud technologies, with a focus on ensuring system stability, scalability, and observability. This position offers the flexibility of remote work, with a preference for candidates located within a few hours of our New Haven office, to promote effective collaboration and support.
Qualifications
- Bachelor’s degree in computer science, engineering, or equivalent practical experience.
- Three or more years of experience in DevOps, infrastructure, systems engineering, or a related field.
- Strong Linux fundamentals and troubleshooting skills in distributed environments.
- Hands‑on experience with AWS and cloud‑based infrastructure.
- Experience working with CI/CD systems such as GitHub Actions.
- Proficiency with infrastructure‑as‑code tools like Terraform.
- Familiarity with containerization using Docker and deploying containerized applications.
- Knowledge of Kubernetes concepts and container orchestration environments.
- Experience with monitoring and observability tools such as Zabbix, CloudWatch, OpenSearch, and Grafana.
- Understanding of relational databases including PostgreSQL, MySQL, and Amazon Aurora, with basic performance troubleshooting skills.
- Scripting experience in Python, Bash, or similar languages to support automation and debugging.
- Solid understanding of networking fundamentals including DNS, routing, and firewalls.
- Strong problem-solving skills with the ability to troubleshoot across multiple systems.
- Ability to work effectively in a fast-paced environment with evolving priorities.
- Desire to learn and grow in areas such as Kubernetes, observability, infrastructure automation, and platform reliability.
Responsibilities
- Develop and maintain logging, metrics, dashboards, and alerting solutions to enhance system observability.
- Utilize tools such as Zabbix, OpenSearch, CloudWatch, and Grafana to improve system visibility, monitoring coverage, and operational reliability across hybrid cloud and hardware‑integrated platforms.
- Investigate and resolve issues related to infrastructure, Kubernetes workloads, CI/CD pipelines, and cloud services.
- Support the operation of Kubernetes platforms, including deploying workloads, troubleshooting issues, and improving platform reliability both on‑premises and in cloud environments.
- Build, maintain, and troubleshoot CI/CD pipelines, primarily using GitHub Actions, to ensure reliable and repeatable deployments.
- Assist with troubleshooting and enhancing the reliability of stateful services and relational database platforms such as PostgreSQL, MySQL, and Amazon Aurora.
- Maintain hybrid infrastructure across on‑premises and AWS environments, ensuring seamless integration and performance.
- Implement and manage infrastructure‑as‑code solutions using Terraform and Ansible to automate provisioning and configuration management.
- Support containerized applications by building, deploying, and debugging Docker‑based workloads.
- Improve alert quality by reducing noise and supporting effective incident response processes.
- Contribute to internal platform development efforts, including development environments and shared infrastructure services.
- Collaborate with cross-functional teams across hardware, software, and infrastructure to resolve issues and enhance system reliability.
- Lead automation initiatives to minimize manual operational work and ensure consistency across environments.
Benefits
- Competitive salary aligned with experience and location.
- Flexible remote work arrangements to promote work-life balance.
- Comprehensive health, dental, and vision insurance plans.
- Retirement savings options with company matching contributions.
- Professional development opportunities, including training and certifications.
- Collaborative and innovative work environment focused on cutting-edge technology.
- Paid time off and holidays to support personal and family needs.