Jobs · Engineering

Senior Site Reliability Engineer

Knowmadics · United States · 2 days ago
RemoteRemoteEngineeringFull-time

Duties and Responsibilities

  • Serve as a senior day-to-day technical owner for customer deployments, including systems both inside and outside secure facilities or SCIFs.
  • Spend approximately 30%-50% of time supporting customer sites, releases, maintenance, incident response, operator needs, and stakeholder activities; contribute to infrastructure team priorities during the balance of the role.
  • Monitor and support the availability, performance and security of linux and kubernetes platform deployments + troubleshoot using logs, metrics, events, and system telemetry.
  • Lead or coordinate incident triage, recovery, root-cause analysis and reporting across customer and internal stakeholders.
  • Plan, coordinate, and validate upgrades, patches, configuration changes, smoke tests, acceptance checks, and rollback procedures.
  • Create and maintain runbooks, standard operating procedures, troubleshooting procedures, site configuration baselines, and administrator documentation.
  • Prepare trip plans, access and logistics requirements, agendas, pre-briefs, onsite briefings and debriefings, trip reports, status updates.
  • Gather operational requirements and field feedback, then work with product and engineering leads to implement new platform/infrastructure features.
  • Design, implement, review, test, and maintain reusable IaC for development, QA, RC, and production cloud platform deployments.
  • Design and maintain IaC and deployment automation for disconnected and air-gapped customer tenancies, including environment-specific configuration, offline artifacts, installation and upgrades.
  • Maintain infrastructure delivery and upgrade processes and release artifacts, including versioning, artifact integrity, vulnerability remediation and secrets management.
  • Perform actions to ensure applicable DoD/DoW platform/software infrastructure security and compliance, support platform security hardening, risk assessment, and remediation.
  • Participate in architecture, release, change control, risk, and post-incident reviews, establish operational standards and provide technical guidance or mentorship without requiring formal direct reports.

Qualifications

  • Demonstrated senior-level experience in site reliability engineering, platform or infrastructure engineering, systems administration, field engineering, or a closely related role, including independent ownership of production systems and customer-facing technical work.
  • Advanced linux system administration experience, familiarity with RHEL administration, hardening, patching, storage, networking, service management, and system troubleshooting is preferred.
  • Hands-on experience operating and troubleshooting kubernetes clusters and supporting cluster lifecycle management, upgrades, workload health, networking, storage, and access controls.
  • Strong experience with infrastructure-as-code and configuration management tools such as Terraform and Ansible, including reusable modules, environment-specific configuration, and state management.
  • Experience designing or operating cloud-native infrastructure across development, QA, RC, and production environments, including compute, networking, identity, storage, secrets, observability, and release workflows.
  • Experience deploying and supporting disconnected or air-gapped systems, including offline artifacts, container images or registries, environment configuration, installation, upgrades, and rollback.
  • Experience establishing and using operational monitoring, alerting, dashboards, logs, and incident-management practices, including root-cause analysis and remediation.
  • Experience with DISA Security Technical Implementation Guides and with DoD/DoW security or compliance regimes, including system hardening, vulnerability remediation, and audit or control evidence.
  • Proven ability to communicate with technical and non-technical stakeholders, lead briefings and debriefings, prepare clear technical reports, and represent engineering interests in access-limited environments.
  • Ability to work autonomously, prioritize competing site and infrastructure needs, identify and communicate risk, and coordinate work across Product, Engineering, customer, and collaborator teams.
  • Must hold an active U.S. Security Clearance at Secret or high level - U.S. Citizenship required.

Similar jobs