Site Reliability Engineer
Southern Finger Lakes · Corning, NY · 1 wk ago
Engineering$70–$75/hrFull-time
Contract rate: $70–$75 per hour. Duration: 6 months with possible extension.
About the role
We are looking for an experienced Site Reliability Engineer to strengthen our team’s platform engineering and operational capabilities. You will play a key role in supporting Kubernetes infrastructure managed through Rancher, improving system reliability and automation, and advancing infrastructure-as-code and GitOps practices across our environment.
Responsibilities
- Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency.
- Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher.
- Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support.
- Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches.
- Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
- Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions.
- Identify opportunities to improve platform resilience, observability, security, and maintainability through automation and modern SRE practices.
Requirements
- 5+ years of professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
- Hands-on experience operating and supporting Kubernetes platforms in production environments.
- Strong experience managing Kubernetes clusters in both on-premises and cloud-based environments.
- Strong Linux systems administration skills, including troubleshooting, scripting, networking, and system performance analysis.
- Experience with Rancher for Kubernetes cluster management and platform operations.
- Experience implementing infrastructure-as-code solutions for platform provisioning and lifecycle management.
- Demonstrated success working in Agile teams (Scrum, Kanban).
Skills
- Kubernetes: Cluster operations, upgrades, networking, storage, troubleshooting, and workload support.
- Platform Management: Rancher or similar Kubernetes management platforms.
- Linux: Advanced administration of Linux/Unix systems.
- Infrastructure as Code: Strong IaC experience; Cluster API (CAPI) preferred.
- GitOps/CI-CD: ArgoCD, Git version control, and deployment automation practices.
- Scripting/Automation: Bash, Python, or similar scripting languages for automation and operational tooling.
Preferred Experience
- Hybrid infrastructure spanning on-premises and public cloud platforms (AWS, Azure, GCP).
- Kubernetes ecosystem tooling for observability, logging, monitoring, and alerting.
- Security best practices for Kubernetes and Linux platforms.
- Supporting scientific research environments, high-performance computing, or computational science workflows.
- CI/CD pipeline development and platform automation patterns.
Qualifications
BS in Computer Science, Software Engineering, Information Technology, or related field preferred; or equivalent professional experience.
About Us
US Tech Solutions is a global staff augmentation firm providing a wide range of talent on-demand and total workforce solutions. To learn more, visit www.ustechsolutions.com.