Senior Site Reliability Engineer - TS/SCI (req-276)
About the role
Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are top-tier program and project management, data analytics, and audit services, backed by an integrated approach to operational excellence. We are looking for a dynamic Senior Site Reliability Engineer (SRE) with a Top Secret/SCI clearance to join our team and help accelerate our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions.
The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will ensure the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure.
Responsibilities
- Monitor and manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on Kubernetes.
- Kubernetes Management: Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performance.
- Containerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environments.
- Cloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.).
- Monitoring & Incident Response: Set up monitoring solutions, define alerts, and manage the incident response process for any issues related to Jenkins or Kubernetes clusters.
- Automate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalent.
- Collaborate Across Teams: Work closely with development, services, and operations teams to ensure seamless integration between application development, deployment, and infrastructure.
- Security & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations, including role-based access, encryption, and automated vulnerability scanning.
Requirements
- Active TOP SECRET/SCI clearance or higher is required.
- Bachelor's degree in Computer Science or related field.
- 5+ years of experience working with on-premise and off-premise cloud environments.
- May require occasional travel (up to 25%).
- Experience with AWS and/or Azure.
- Hands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, Terraform, Helm, PostgreSQL, or similar technologies.
- Ability to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScript.
- Experience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn).
- Proactive approach to identifying problems, performance bottlenecks, and areas for improvement.
- Ability to lead and work independently in an Agile/Scrum environment.
- Real passion for developing team-oriented solutions to complex engineering problems.
- Thrive in an autonomous, empowering, and exciting environment.
- Great verbal and written communication skills to collaborate multi-functionally and improve scalability.
- Interest in committing to a fun, friendly, expansive, and intellectually stimulating environment.
Desired Skills
- Hands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud Services.
- Experience with deep learning, natural language processing, computer vision, or reinforcement learning.
- Ability to convey highly technical concepts and information in written form to technical and non-technical audiences.
- Ability to work on multiple concurrent projects.
- Strong self-motivation and the ability to work with minimal supervision.
- Team-oriented individual, energetic, result & delivery oriented, with a keen interest in quality and the ability to meet deadlines.
Benefits
- Performance Bonuses
- Medical Insurance
- Dental Insurance
- Vision Insurance
- 401(k) Plan (Traditional and ROTH)
- Life Insurance (Basic, Voluntary & AD&D)
- Paid Time Off
- 11 Federal Holidays
- Parental Leave
- Commute Benefits
- Short Term & Long Term Disability
- Training & Development
- Wellness Program
- Community Outreach Initiatives
Pay
The annual salary range for this role is $165,000 - $200,000. Please note that the salary information provided is a general guideline. CATHEXIS considers various factors in its final offer, including location, qualifications, experience, and skills.