Jobs · Engineering · Texas

Senior Site Reliability Engineer

O2 Technologies,Inc · Dallas, TX · Yesterday
On-siteEngineering$75–$80/hrContract

Location: Irving, TX (Hybrid 2-3 days a week on site)
Contract length: 1 year (renewable)
Pay rate: $75-80/hr (c2c)

About the role

We are seeking a Senior Kubernetes-focused Site Reliability Engineer with strong cloud automation and software engineering skills who can leverage AI and large-language models to automate operations and improve platform reliability at scale.

Responsibilities

  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
  • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
  • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Requirements

  • 5+ years of hands-on Kubernetes platform engineering with GKE and Rancher RKE2, including multi-cluster management, troubleshooting, and performance optimization.
  • 5+ years of advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
  • Strong experience in GCP, Terraform, Helm, GitHub, and CI/CD pipelines for production-grade automation.
  • Experience with observability and monitoring tools: Splunk, Grafana, Datadog, AppDynamics.
  • Experience with API and microservices engineering: Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
  • Experience applying LLMs (Gemini, Llama, Mistral, Qwen) for alert analysis, incident triage, automation, and operational workflows.
  • Bachelor’s degree in Computer Science, Software Engineering, Electrical Engineering, Information Technology, or Computer Engineering.
  • Preferred: Master’s or PhD in Computer Science, Software Engineering, Data Science, or related field.
  • Industry experience in healthcare, cloud platform engineering, site reliability engineering, or AI/machine learning operations.

Skills

  • Site Reliability Engineering (SRE): reliability, availability, incident management, SLO/SLI monitoring, operational excellence.
  • Kubernetes platform engineering: multi-cluster and multi-datacenter management, disaster recovery, active-active deployments.
  • Cloud automation and infrastructure-as-code (IaC).
  • AI-driven operations (AIOps) with LLMs.
  • Observability and monitoring.
  • API and microservices reliability.
  • Cross-functional reliability partnerships.

Similar jobs