AI Ops / DevOps Engineer
CareerZen · Atlanta, GA · 3 days ago
Engineering$60–$70/hrFull-time
About the role
The Senior AI Ops / DevOps Engineer will architect, build, and manage next-generation AI-driven CI/CD and cloud operations ecosystems. This role will go beyond traditional DevOps automation by integrating LLM agents, Model Context Protocol servers, intelligent observability, and secure AI-assisted workflows into the software delivery lifecycle.
Responsibilities
- Architect, build, and manage AI-enabled CI/CD pipelines that improve developer productivity, code quality, release reliability, and deployment speed.
- Design and deploy production-grade Model Context Protocol clients and servers to securely connect enterprise LLMs with engineering tools, repositories, cloud infrastructure, and observability platforms.
- Develop custom MCP servers using Python, TypeScript, Node.js, or JavaScript to expose logs, infrastructure metrics, deployment data, and internal tools to authorized AI agents.
- Integrate LLM agents into developer workflows to support automated code review, vulnerability detection, test generation, release validation, and infrastructure recommendations.
- Create safe autonomous remediation workflows for log analysis, incident triage, root-cause analysis, and infrastructure issue resolution.
- Build and maintain robust CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, ArgoCD, Jenkins, or similar tools.
- Implement ChatOps 2.0 capabilities that allow engineers to interact with deployment pipelines, cloud environments, logs, and operational workflows using secure conversational interfaces.
- Manage containerized workloads using Docker and Kubernetes platforms such as AWS EKS, Azure AKS, or Google GKE.
- Integrate AI-driven observability workflows with platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK.
- Implement AI safety controls including role-based access control, least-privilege execution, human-in-the-loop approvals, audit logging, rollback mechanisms, and secure tool access.
- Partner with software engineering, DevOps, SRE, security, platform, and data/AI teams to identify opportunities for intelligent automation.
- Create reusable automation frameworks, runbooks, dashboards, documentation, and enablement materials for engineering teams.
- Drive an "automate everything" culture by reducing manual toil and improving operational efficiency across cloud and software delivery processes.
Requirements
- Minimum 7+ years of experience in DevOps, Cloud Engineering, SRE, Platform Engineering, or Infrastructure Automation.
- Minimum 4+ years of hands-on experience designing and managing CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD, or similar platforms.
- Minimum 3+ years of experience managing scalable cloud environments in AWS, Azure, or GCP, with strong preference for AWS.
- Minimum 3+ years of Strong hands-on experience with Kubernetes, Docker, and production container orchestration platforms such as EKS, AKS, or GKE.
- Advanced proficiency with Infrastructure as Code tools such as Terraform, OpenTofu, Terragrunt, Pulumi, CloudFormation, or Crossplane.
- Minimum 3+ years of Strong programming and scripting experience using Python, TypeScript, JavaScript, Bash, or Go.
- Minimum 3+ years of Practical experience working with LLM APIs such as OpenAI, Anthropic, or similar enterprise AI platforms.
- Minimum 3+ years of Experience with AI orchestration or agentic frameworks such as LangChain, CrewAI, LlamaIndex, or similar tools.
- Minimum 3+ years of Strong understanding of the Model Context Protocol ecosystem and experience designing or integrating MCP clients and servers.
- Minimum 3+ years of Experience integrating DevSecOps controls into CI/CD pipelines, including SAST, DAST, dependency scanning, container scanning, secrets scanning, and vulnerability management.
- Minimum 3+ years of Strong knowledge of secret management and security tooling such as HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or similar platforms.
- Minimum 3+ years of Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK.
- Familiarity with security and compliance frameworks such as SOC2, ISO27001, or enterprise audit control environments.
- Ability to troubleshoot complex pipeline, infrastructure, deployment, and production issues across cloud-native environments.
Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent work experience.
Skills
- Strong "automate everything" mindset with a passion for reducing repetitive manual tasks and operational toil.
- Security-first approach with practical skepticism of autonomous AI actions and a focus on validation, boundaries, approvals, and rollback.
- Collaborative educator who can help upskill engineering teams on AI-assisted delivery, secure automation, and modern DevOps practices.
- Ownership mindset with the ability to design solutions, implement them hands-on, and support them in production.
Benefits
- Starting hourly range for this remote role is ($60-$70/hour).
- This position may also be eligible for incentive compensation based on individual and/or company performance.
- This position is eligible for company benefits that will depend on the nature of the role offered.
Pay
This position is eligible for a starting hourly range of ($60-$70/hour).
Schedule
This position requires 3 days in office per the client/project requirement.