Jobs · Information Technology · Virginia

Senior Site Reliability Engineer - TS/SCI (req-276)

CATHEXIS · Tysons Corner, VA · 4 wk ago
On-siteInformation Technology$165k–$200k/yrFull-time

About the role

Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are top-tier program and project management, data analytics, and audit services, backed by an integrated approach to operational excellence. We are looking for a dynamic Senior Site Reliability Engineer (SRE) with a Top Secret/SCI clearance to join our team and help accelerate our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions.

The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will ensure the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure.

Responsibilities

  • Monitor and manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on Kubernetes.
  • Kubernetes Management: Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performance.
  • Containerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environments.
  • Cloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.).
  • Monitoring & Incident Response: Set up monitoring solutions, define alerts, and manage the incident response process for any issues related to Jenkins or Kubernetes clusters.
  • Automate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalent.
  • Collaborate Across Teams: Work closely with development, services, and operations teams to ensure seamless integration between application development, deployment, and infrastructure.
  • Security & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations, including role-based access, encryption, and automated vulnerability scanning.

Requirements

  • Active TOP SECRET/SCI clearance or higher is required.
  • Bachelor's degree in Computer Science or related field.
  • 5+ years of experience working with on-premise and off-premise cloud environments.
  • May require occasional travel (up to 25%).
  • Experience with AWS and/or Azure.
  • Hands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, Terraform, Helm, PostgreSQL, or similar technologies.
  • Ability to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScript.
  • Experience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn).
  • Proactive approach to identifying problems, performance bottlenecks, and areas for improvement.
  • Ability to lead and work independently in an Agile/Scrum environment.
  • Real passion for developing team-oriented solutions to complex engineering problems.
  • Thrive in an autonomous, empowering, and exciting environment.
  • Great verbal and written communication skills to collaborate multi-functionally and improve scalability.
  • Interest in committing to a fun, friendly, expansive, and intellectually stimulating environment.

Desired Skills

  • Hands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud Services.
  • Experience with deep learning, natural language processing, computer vision, or reinforcement learning.
  • Ability to convey highly technical concepts and information in written form to technical and non-technical audiences.
  • Ability to work on multiple concurrent projects.
  • Strong self-motivation and the ability to work with minimal supervision.
  • Team-oriented individual, energetic, result & delivery oriented, with a keen interest in quality and the ability to meet deadlines.

Benefits

  • Performance Bonuses
  • Medical Insurance
  • Dental Insurance
  • Vision Insurance
  • 401(k) Plan (Traditional and ROTH)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off
  • 11 Federal Holidays
  • Parental Leave
  • Commute Benefits
  • Short Term & Long Term Disability
  • Training & Development
  • Wellness Program
  • Community Outreach Initiatives

Pay

The annual salary range for this role is $165,000 - $200,000. Please note that the salary information provided is a general guideline. CATHEXIS considers various factors in its final offer, including location, qualifications, experience, and skills.

Similar jobs