Jobs · OTHR · New York

Systems Reliability Engineer - Remote (EST)

The Fountain Group · Corning, NY · 1 wk ago
On-siteOTHR$70–$75/hrContract

Sorry, NO C2C or sponsorship provided. This is a 6-month remote assignment (supporting EST) with possible extension based on budget and performance. Candidates can live anywhere in the US but must be able to work 8am–5pm or 9am–6pm EST.

About the Role

This position will help maintain, enhance, and evolve the Kubernetes platforms that enable scientific and engineering teams to deploy, operate, and scale critical applications across both on-premises and cloud environments. You will strengthen the team’s platform engineering and operational capabilities by supporting Kubernetes infrastructure managed through Rancher, improving system reliability and automation, and advancing infrastructure-as-code and GitOps practices.

Responsibilities

  • Platform Operations: Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency.
  • Cluster Management: Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher.
  • Linux Systems Administration: Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support.
  • Infrastructure as Code: Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches.
  • GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
  • Collaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions.
  • Continuous Improvement: Identify opportunities to improve platform resilience, observability, and security.

Requirements

  • Minimum of 5 years professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
  • 5 years strong system administration with Linux.
  • Rancher for Kubernetes experience is required.
  • Hands-on experience operating and supporting Kubernetes platforms in production environments.
  • Strong experience managing Kubernetes clusters in both on-premises and cloud-based environments.
  • Strong Linux systems administration skills, including troubleshooting, scripting, networking, and system performance analysis.
  • Experience implementing infrastructure-as-code solutions for platform provisioning and lifecycle management; Cluster API (CAPI) preferred.
  • Demonstrated success working in Agile teams (Scrum, Kanban).
  • Experience with GitOps/CI-CD: ArgoCD, Git version control, and deployment automation practices.
  • Scripting/Automation: Bash, Python, or similar scripting languages for automation and operational tooling.

Preferred Qualifications

  • Experience with hybrid infrastructure spanning on-premises and public cloud platforms (AWS, Azure, GCP).
  • Experience with Kubernetes ecosystem tooling for observability, logging, monitoring, and alerting.
  • Familiarity with security best practices for Kubernetes and Linux platforms.
  • Experience supporting scientific research environments, high-performance computing, or computational science workflows.
  • Knowledge of CI/CD pipeline development and platform automation patterns.

Pay

$70–$75/hr to start.

Schedule

6-month assignment, possible extension based on budget and performance. Must be available 8am–5pm or 9am–6pm EST.

Similar jobs