Cloud Platform Engineer (Site Reliability)
About the role
We are seeking a Cloud Platform Engineer (Site Reliability) to join our team Nexus, a teammate company working with NASA. In this role, you will contribute to NASA’s deep space exploration missions, including Orion, Lunar Gateway, and Artemis, while developing cloud-native platform services that benefit people on Earth. This position offers the opportunity to work with a passionate and diverse team, shaping systems that inspire global innovation.
Responsibilities
- Develop new cloud-native platform services spanning all three major cloud environments
- Establish and promote best practices for cloud-native application development within the organization
- Administer NASA cloud networks and manage deployment requests for COTS and cloud-native applications
- Write quality code and provide engaged, constructive code reviews for peers
- Work with managed Kubernetes offerings across all three major cloud providers
- Integrate cloud-managed AI and data services with bespoke and open-source Kubernetes applications
- Identify opportunities to abstract project requirements and develop enterprise-grade, multi-tenant platform services
- Collaborate with NASA security and compliance teams to ensure adherence to industry best practices and regulatory requirements
- Work directly with NASA human spaceflight missions such as Orion, Lunar Gateway, and Artemis
- Perform other duties as required
Requirements
- Bachelor’s degree or equivalent certification in a related field
- Typically 10 years of experience in a related area
- Strong experience with Kubernetes in production
- Proficiency in managing and using GitLab
- Hands-on experience with CI/CD pipeline tools
- Experience with observability monitoring tools such as Grafana and Superset
- Proficiency with Infrastructure-as-Code using Terraform or open-source alternatives (e.g., OpenTofu)
- Extensive Linux experience (familiarity with Windows preferred but not required)
- Expertise in at least one programming language (Go and Python preferred)
- Experience with Python and SQL (R is a plus)
- Working understanding of Machine Learning Model Lifecycle management (preferred)
Preferred Qualifications
- Solution-oriented mindset with enthusiasm for feedback
- Strong communication skills, including clear status updates and collaborative decision-making
- Self-motivated and self-managing with excellent organizational skills
- Experience architecting and building cloud-native applications
- Demonstrable success in owning the full project lifecycle
- Instinct for identifying abstraction opportunities
- Ability to balance individual project tasks with broader platform and company goals
Benefits
- Excellent personal and professional career growth opportunities
- 9/80 work schedule (every other Friday off), where applicable
- Onsite cafeteria offering breakfast and lunch
- Health, dental, and vision insurance
- Paid time off and holidays
- Retirement benefits, including 401(k) matching
- Educational reimbursement
- Parental leave
- Employee stock purchase plan
- Tax-saving options
- Disability and life insurance
- Pet insurance
Note: Benefits may vary based on employment type, location, and applicable agreements. Positions governed by a Collective Bargaining Agreement (CBA), the McNamara-O'Hara Service Contract Act (SCA), or other employment contracts may include different provisions.
Additional Information
Proof of U.S. Citizenship or U.S. Permanent Residency may be required. Must be able to complete a U.S. government background investigation. This position is posted at multiple levels, and the company reserves the right to consider candidates at any level based on experience, requirements, and business needs.