Cloud Platform Engineer
About the role
At Farmers, we deliver peace of mind when it matters most through a results-driven, high-performance culture that thrives on creativity, accountability, and bold solutions. The Cloud Platform Operations team is responsible for the analysis, design, implementation, and operational support of multi-cloud infrastructure across GCP (Google Cloud Platform), Azure (Microsoft cloud), and AWS. This role ensures platform availability, scalability, and compliance while driving AIOps/FinOps and autonomous infrastructure management. It focuses on leveraging Multi-Cloud AI offerings to drive operational reliability, predictive monitoring, and continuous optimization to maintain high availability and resilience in production environments.
Responsibilities
- Ensure availability, scalability, reliability, and security of cloud platforms and services.
- Design, deploy, and govern AI-powered agents (e.g., using Azure Copilot/AWS Bedrock) to achieve autonomous self-healing capabilities and automated resource management.
- Handle regular operational requests with hands-on experience using Terraform for EC2 changes, S3 updates, user access management, and managed services like SageMaker, Bedrock, Storage Gateway, RDS, and Transfer Family.
- Supervise and refine AI-generated Infrastructure-as-Code (IaC) (Terraform/Ansible) for developing and maintaining complex and scalable Terraform/Ansible/CloudBees (Jenkins) automation pipelines to provision, deploy, patch, and manage cloud infrastructure.
- Implement AI-based automation solutions for Cloud Operations to monitor performance, scalability, and respond to incidents and operational issues autonomously.
- Implement GenAI tools to perform real-time Root Cause Analysis (RCA), correlate complex event data (logs, metrics), and auto-generate runbooks and incident summaries.
- Develop and train predictive ML models to analyze historical telemetry and forecast potential system outages or performance bottlenecks, and configure proactive monitoring and alerting for critical services.
- Manage security and compliance by utilizing AI agents to detect configuration drift and auto-generate compliant updates for IAM, network, and security policies.
- Remediate vulnerabilities, configure notifications, support audits, and maintain certifications and governance standards.
- Collaborate with application, architecture, AIOps, FinOps, and security teams to ensure production readiness.
- Lead proof-of-concepts, implement new cloud services, and adapt quickly to new cloud releases and features to enhance operational capabilities.
- Understand middleware components to provide end-to-end production support and troubleshoot complex operational scenarios.
- Work with application teams, analyzing logs and data, opening service requests, working with vendors and partners, and driving problem resolution.
- Review architectural designs for applications to ensure reliable and performant design patterns are implemented.
- Deploy applications, workloads, and data to the cloud environment, often involving migration from on-premises infrastructure or other cloud providers.
- Work with Finance and procurement teams to implement cost optimization strategies based on changing workload patterns, business requirements, and new offerings from cloud providers.
Requirements
- 5 years of experience in public cloud operations (AWS, Azure, GCP), with a strong focus on AIOps integration.
- Deep, demonstrable expertise designing and operationalizing solutions leveraging AWS Bedrock/Agent Frameworks and Azure Copilot for Cloud Operations.
- Expertise in Infrastructure as Code (Terraform, CloudFormation), Ansible, and CI/CD pipelines, including supervising AI-generated infrastructure artifacts.
- Proficiency in scripting languages (Python, Bash).
- Expertise in integrating observability platforms (Dynatrace, Prometheus) into AI/ML platforms for predictive analysis and anomaly detection.
- Understanding of Site Reliability Engineering (SRE) and operational reliability principles.
- Experience with monitoring tools (CloudWatch, Prometheus, Dynatrace, Azure Monitor) and ServiceNow.
- Ability to adapt to new cloud releases and emerging technologies.
- Excellent troubleshooting, problem-solving, and communication skills.
- Hands-on experience with JBoss, Tomcat, IBM WAS, Oracle, DB2, MSSQL, and Postgres is a plus.
Qualifications
- Bachelor’s degree in computer science or other technical discipline.
- Technical Certifications are a plus.
Physical Requirements
- This role includes normal and customary distractions, noise, and interruptions.
- Sits or stands for extended periods of time, up to a full work shift.
- Occasionally reaches overhead and below the knees, including bending, twisting, pulling, and stooping.
- Occasionally moves, lifts, carries, and places objects and supplies weighing 0-10 pounds without assistance.
- Listens to, interprets, and differentiates auditory information (e.g., others speaking) at normal speaking levels with or without correction.
- Visually verifies and reads information, and locates material, resources, and other objects.
- Ability to continuously operate a computer for extended periods of time, up to a full work shift.
- Physical dexterity sufficient to use hands, arms, and shoulders repetitively to operate keyboard and other office equipment up to a full work shift.
Pay
- CA Only: $102,450 - $174,240
- CO Only: $96,075 - $150,260
- HI/IL/MN/VT Only: $96,075 - $160,710
- MA Only: $96,075 - $160,710
- MD Only: $96,075 - $160,710
- DC/NJ/NY/OH Only: $96,075 - $174,240
- Albany County, NY/Cleveland, OH: $102,450 - $150,260
- WA Only: $96,075 - $182,625
- Bonus Opportunity (based on Company and Individual Performance)
Benefits
- 401(k)
- Medical
- Dental
- Vision
- Health Savings and Flexible Spending Accounts
- Life Insurance
- Paid Time Off
- Paid Parental Leave
- Tuition Assistance
For more information, review “What we offer” on Farmers Careers.