Senior Site Reliability Engineer
O2 Technologies,Inc · Dallas, TX · Yesterday
On-siteEngineering$75–$80/hrContract
Location: Irving, TX (Hybrid 2-3 days a week on site)
Contract length: 1 year (renewable)
Pay rate: $75-80/hr (c2c)
About the role
We are seeking a Senior Kubernetes-focused Site Reliability Engineer with strong cloud automation and software engineering skills who can leverage AI and large-language models to automate operations and improve platform reliability at scale.
Responsibilities
- Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
- Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
- Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
- Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
- Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
- Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
- Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Requirements
- 5+ years of hands-on Kubernetes platform engineering with GKE and Rancher RKE2, including multi-cluster management, troubleshooting, and performance optimization.
- 5+ years of advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
- Strong experience in GCP, Terraform, Helm, GitHub, and CI/CD pipelines for production-grade automation.
- Experience with observability and monitoring tools: Splunk, Grafana, Datadog, AppDynamics.
- Experience with API and microservices engineering: Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
- Experience applying LLMs (Gemini, Llama, Mistral, Qwen) for alert analysis, incident triage, automation, and operational workflows.
- Bachelor’s degree in Computer Science, Software Engineering, Electrical Engineering, Information Technology, or Computer Engineering.
- Preferred: Master’s or PhD in Computer Science, Software Engineering, Data Science, or related field.
- Industry experience in healthcare, cloud platform engineering, site reliability engineering, or AI/machine learning operations.
Skills
- Site Reliability Engineering (SRE): reliability, availability, incident management, SLO/SLI monitoring, operational excellence.
- Kubernetes platform engineering: multi-cluster and multi-datacenter management, disaster recovery, active-active deployments.
- Cloud automation and infrastructure-as-code (IaC).
- AI-driven operations (AIOps) with LLMs.
- Observability and monitoring.
- API and microservices reliability.
- Cross-functional reliability partnerships.