Jobs · Engineering

Site Reliability Engineer - FedRAMP

Jobgether · United States · 1 mo ago
RemoteRemoteEngineering$152k–$253k/yrFull-time

Accountabilities

  • Maintain and improve the operational foundation of a secure, cloud-based platform.
  • Gain deep understanding of platform workloads, dependencies, and operational workflows through documentation, code analysis, and collaboration with subject matter experts.
  • Create and maintain operational documentation, including runbooks, incident guides, onboarding resources, and knowledge-sharing materials.
  • Participate in incident response activities, including investigation, mitigation, root cause analysis, and post-incident improvements.
  • Support the implementation and maintenance of reliability practices, including SLIs, SLOs, error budgets, and availability improvements.
  • Improve system observability by developing monitoring, alerting, dashboards, and instrumentation strategies.
  • Reduce operational complexity through automation, tooling improvements, and elimination of repetitive tasks.
  • Support infrastructure delivery through infrastructure-as-code, CI/CD pipelines, deployment workflows, and configuration management.
  • Contribute to secure and compliant infrastructure changes within regulated environments.
  • Collaborate with engineering, security, compliance, and operations teams to improve reliability, communicate risks, and resolve technical challenges.
  • Participate in on-call rotations and help ensure platform stability and resilience.

Requirements

  • Strong software engineering fundamentals, cloud infrastructure experience, and the ability to operate effectively in regulated environments.
  • Comfortable investigating complex systems, improving reliability practices, and collaborating across technical teams.
  • 3+ years of experience in software engineering, including at least 1 year working in Site Reliability Engineering, Platform Engineering, or DevOps roles supporting cloud-hosted services.
  • Experience with cloud infrastructure platforms such as Azure or similar cloud providers.
  • Familiarity with compliance-focused environments such as government, FedRAMP, CMMC, financial services, or healthcare industries.
  • Ability to understand and troubleshoot application code to investigate system behavior independently.
  • Experience with observability and monitoring tools such as Prometheus, Grafana, OpenTelemetry, or ELK stack.
  • Hands-on experience with infrastructure-as-code tools such as Terraform, Terragrunt, or Pulumi.
  • Experience with container orchestration platforms, particularly Kubernetes.
  • Experience managing CI/CD workflows using tools such as GitHub Actions, Azure DevOps, GitLab CI, or ArgoCD.
  • Strong programming skills in languages such as TypeScript, JavaScript, Go, Java, C#, or similar.
  • Understanding of distributed systems concepts, networking fundamentals, and cloud reliability principles.
  • Strong written and verbal communication skills with the ability to explain technical concepts clearly.
  • Experience with government or sovereign cloud environments, SaaS platforms, multi-tenant systems, resilience testing, or chaos engineering is a plus.
  • Familiarity with AI-assisted development workflows and LLM-powered tools for automation, documentation, or engineering productivity is beneficial.

Benefits

  • Competitive compensation package with geographic-based salary ranges.
  • Unlimited paid time off and 12 paid holidays, including dedicated company wellness days.
  • Paid volunteer time and community engagement opportunities.
  • Paid parental leave programs.
  • Medical, dental, and vision coverage starting from the first day.
  • Mental health support, therapy resources, and digital wellness programs.
  • 401(k) retirement plan with company matching contributions.
  • Fertility, adoption, and surrogacy support.
  • Virtual veterinary care benefits.
  • Legal services, identity protection, and supplemental insurance options.
  • Healthcare, dependent care, and commuting spending accounts.
  • Access to professional development resources, learning platforms, mentoring, workshops, and training opportunities.
  • Opportunity to work remotely while collaborating with global engineering teams.
  • Exposure to impactful cloud reliability projects within secure and regulated environments.

Similar jobs

Build Reliability Engineer

Varda Space IndustriesEl Segundo, CA· 2 mo ago
Engineering$109k–$145k/yrapply on job-boards.greenhouse.io

Reliability Engineer

I-care Reliability Inc.Pennsylvania, United States· 1 mo ago
Engineeringapply on icare-usa.cvw.io

Reliability Engineer

Amphenol Communications SolutionsRaleigh, NC· 3 wk ago
Engineeringapply on amphenol-tcs.acquiretm.com

Reliability Engineer

Amphenol Communications SolutionsYocumtown, PA· 3 wk ago
Engineeringapply on amphenol-tcs.acquiretm.com