Staff DevOps Engineer
Northrop Grumman · Aurora, CO · 1 wk ago
On-siteEngineering$177k–$266k/yrFull-time
About the role
Northrop Grumman’s Space Sector is seeking a Staff DevOps Engineer to support on-premises mission systems by operating, maintaining, and incrementally enhancing existing DevOps and platform solutions.
Responsibilities
- Operate and maintain existing GitLab CI/CD infrastructure including patching and permissions.
- Advises on updates to jobs, runners, variables, and approvals to improve stability and throughput.
- Administers Nexus artifact repositories including patching, backups, access management, repository configuration, cleanup/retention policies, and routine health checks.
- Maintains Apache NiFi infrastructure including patching, troubleshooting, performance tuning, and implementing changes defined by senior engineers.
- Supports installation and day-to-day operations of Kubernetes on-prem clusters, including patching, storage management, deployments, rollouts/rollbacks, configuration updates, and resource adjustments.
- Maintains and extends monitoring and alerting using Grafana and Prometheus for on-prem infrastructure and applications, including patching, access control, refining dashboards, alert rules, and service onboarding within existing observability patterns.
- Implements incremental improvements to existing DevOps tooling and workflows (e.g., adding tests, improving pipeline stages, automating manual steps) consistent with established designs and security baselines.
- Performs routine incident response and problem resolution for build, deployment, and runtime issues across CI/CD, Kubernetes, Nexus, NiFi, and supporting services.
- Applies existing security and compliance controls across the DevOps toolchain: secrets management, access controls, certificate and credential rotation, image scanning, and log retention in accordance with program requirements.
- Follows formal change management processes, including documenting changes, participating in reviews, and updating runbooks, SOPs, and technical documentation.
- Collaborates with senior engineers, system administrators, and developers to implement improvements and operational changes defined by leads and architects.
Requirements
- Bachelor’s in STEM degree with 12 years of professional experience; Master’s in STEM degree with 10 years of professional experience; PhD with 7 years of experience.
- Will consider an additional 4+ years of experience in lieu of degree.
- Requires an active Top-Secret/Sensitive Compartmented Information (SCI) clearance at time of application.
- Must have the ability to obtain and maintain U.S. Government Full Scope Polygraph.
- Relevant DevOps, site reliability, or systems engineering experience.
- Demonstrated experience maintaining GitLab CI/CD in an on-prem environment.
- Practical experience administering Nexus repositories.
- Experience supporting Apache NiFi infrastructure.
- Working knowledge of on-prem Kubernetes operations.
- Experience using Grafana and Prometheus in an on-prem context.
- Proficiency with Linux systems administration fundamentals and at least one scripting language (e.g., Bash, Python) for automation and operational tasks.
- Familiarity with operating in secure, classified, or highly regulated on-prem environments and adherence to associated processes (security, configuration control, and auditing).
- Strong analytical and troubleshooting skills; ability to methodically diagnose issues across multiple interconnected services.
Qualifications
- Relevant DevOps, site reliability, or systems engineering experience.
- Demonstrated experience maintaining GitLab CI/CD in an on-prem environment.
- Practical experience administering Nexus repositories.
- Experience supporting Apache NiFi infrastructure.
- Working knowledge of on-prem Kubernetes operations.
- Experience using Grafana and Prometheus in an on-prem context.
- Proficiency with Linux systems administration fundamentals and at least one scripting language (e.g., Bash, Python) for automation and operational tasks.
- Familiarity with operating in secure, classified, or highly regulated on-prem environments and adherence to associated processes (security, configuration control, and auditing).
- Strong analytical and troubleshooting skills; ability to methodically diagnose issues across multiple interconnected services.
Skills
- Experience using Helm and Rancher.
- Familiarity with Infrastructure as Code and configuration management tools (e.g., Terraform, Ansible) for maintaining existing on-prem infrastructure definitions and playbooks.
- Experience with log aggregation and analysis platforms in on-prem environments (e.g., ELK/EFK or similar).
- Experience working within formal change control and release management processes.
- Exposure to microservices running on Kubernetes and related operational patterns (service discovery, scaling, health checks).
Benefits
This position is contingent upon the candidate obtaining final clearances and program access(es) within a reasonable period of time as determined by the company.