Senior Site Reliability Engineer
Knowmadics · United States · 2 days ago
RemoteRemoteEngineeringFull-time
Duties and Responsibilities
- Serve as a senior day-to-day technical owner for customer deployments, including systems both inside and outside secure facilities or SCIFs.
- Spend approximately 30%-50% of time supporting customer sites, releases, maintenance, incident response, operator needs, and stakeholder activities; contribute to infrastructure team priorities during the balance of the role.
- Monitor and support the availability, performance and security of linux and kubernetes platform deployments + troubleshoot using logs, metrics, events, and system telemetry.
- Lead or coordinate incident triage, recovery, root-cause analysis and reporting across customer and internal stakeholders.
- Plan, coordinate, and validate upgrades, patches, configuration changes, smoke tests, acceptance checks, and rollback procedures.
- Create and maintain runbooks, standard operating procedures, troubleshooting procedures, site configuration baselines, and administrator documentation.
- Prepare trip plans, access and logistics requirements, agendas, pre-briefs, onsite briefings and debriefings, trip reports, status updates.
- Gather operational requirements and field feedback, then work with product and engineering leads to implement new platform/infrastructure features.
- Design, implement, review, test, and maintain reusable IaC for development, QA, RC, and production cloud platform deployments.
- Design and maintain IaC and deployment automation for disconnected and air-gapped customer tenancies, including environment-specific configuration, offline artifacts, installation and upgrades.
- Maintain infrastructure delivery and upgrade processes and release artifacts, including versioning, artifact integrity, vulnerability remediation and secrets management.
- Perform actions to ensure applicable DoD/DoW platform/software infrastructure security and compliance, support platform security hardening, risk assessment, and remediation.
- Participate in architecture, release, change control, risk, and post-incident reviews, establish operational standards and provide technical guidance or mentorship without requiring formal direct reports.
Qualifications
- Demonstrated senior-level experience in site reliability engineering, platform or infrastructure engineering, systems administration, field engineering, or a closely related role, including independent ownership of production systems and customer-facing technical work.
- Advanced linux system administration experience, familiarity with RHEL administration, hardening, patching, storage, networking, service management, and system troubleshooting is preferred.
- Hands-on experience operating and troubleshooting kubernetes clusters and supporting cluster lifecycle management, upgrades, workload health, networking, storage, and access controls.
- Strong experience with infrastructure-as-code and configuration management tools such as Terraform and Ansible, including reusable modules, environment-specific configuration, and state management.
- Experience designing or operating cloud-native infrastructure across development, QA, RC, and production environments, including compute, networking, identity, storage, secrets, observability, and release workflows.
- Experience deploying and supporting disconnected or air-gapped systems, including offline artifacts, container images or registries, environment configuration, installation, upgrades, and rollback.
- Experience establishing and using operational monitoring, alerting, dashboards, logs, and incident-management practices, including root-cause analysis and remediation.
- Experience with DISA Security Technical Implementation Guides and with DoD/DoW security or compliance regimes, including system hardening, vulnerability remediation, and audit or control evidence.
- Proven ability to communicate with technical and non-technical stakeholders, lead briefings and debriefings, prepare clear technical reports, and represent engineering interests in access-limited environments.
- Ability to work autonomously, prioritize competing site and infrastructure needs, identify and communicate risk, and coordinate work across Product, Engineering, customer, and collaborator teams.
- Must hold an active U.S. Security Clearance at Secret or high level - U.S. Citizenship required.