Site Reliability Engineer - FedRAMP
Jobgether · United States · 1 mo ago
RemoteRemoteEngineering$152k–$253k/yrFull-time
Accountabilities
- Maintain and improve the operational foundation of a secure, cloud-based platform.
- Gain deep understanding of platform workloads, dependencies, and operational workflows through documentation, code analysis, and collaboration with subject matter experts.
- Create and maintain operational documentation, including runbooks, incident guides, onboarding resources, and knowledge-sharing materials.
- Participate in incident response activities, including investigation, mitigation, root cause analysis, and post-incident improvements.
- Support the implementation and maintenance of reliability practices, including SLIs, SLOs, error budgets, and availability improvements.
- Improve system observability by developing monitoring, alerting, dashboards, and instrumentation strategies.
- Reduce operational complexity through automation, tooling improvements, and elimination of repetitive tasks.
- Support infrastructure delivery through infrastructure-as-code, CI/CD pipelines, deployment workflows, and configuration management.
- Contribute to secure and compliant infrastructure changes within regulated environments.
- Collaborate with engineering, security, compliance, and operations teams to improve reliability, communicate risks, and resolve technical challenges.
- Participate in on-call rotations and help ensure platform stability and resilience.
Requirements
- Strong software engineering fundamentals, cloud infrastructure experience, and the ability to operate effectively in regulated environments.
- Comfortable investigating complex systems, improving reliability practices, and collaborating across technical teams.
- 3+ years of experience in software engineering, including at least 1 year working in Site Reliability Engineering, Platform Engineering, or DevOps roles supporting cloud-hosted services.
- Experience with cloud infrastructure platforms such as Azure or similar cloud providers.
- Familiarity with compliance-focused environments such as government, FedRAMP, CMMC, financial services, or healthcare industries.
- Ability to understand and troubleshoot application code to investigate system behavior independently.
- Experience with observability and monitoring tools such as Prometheus, Grafana, OpenTelemetry, or ELK stack.
- Hands-on experience with infrastructure-as-code tools such as Terraform, Terragrunt, or Pulumi.
- Experience with container orchestration platforms, particularly Kubernetes.
- Experience managing CI/CD workflows using tools such as GitHub Actions, Azure DevOps, GitLab CI, or ArgoCD.
- Strong programming skills in languages such as TypeScript, JavaScript, Go, Java, C#, or similar.
- Understanding of distributed systems concepts, networking fundamentals, and cloud reliability principles.
- Strong written and verbal communication skills with the ability to explain technical concepts clearly.
- Experience with government or sovereign cloud environments, SaaS platforms, multi-tenant systems, resilience testing, or chaos engineering is a plus.
- Familiarity with AI-assisted development workflows and LLM-powered tools for automation, documentation, or engineering productivity is beneficial.
Benefits
- Competitive compensation package with geographic-based salary ranges.
- Unlimited paid time off and 12 paid holidays, including dedicated company wellness days.
- Paid volunteer time and community engagement opportunities.
- Paid parental leave programs.
- Medical, dental, and vision coverage starting from the first day.
- Mental health support, therapy resources, and digital wellness programs.
- 401(k) retirement plan with company matching contributions.
- Fertility, adoption, and surrogacy support.
- Virtual veterinary care benefits.
- Legal services, identity protection, and supplemental insurance options.
- Healthcare, dependent care, and commuting spending accounts.
- Access to professional development resources, learning platforms, mentoring, workshops, and training opportunities.
- Opportunity to work remotely while collaborating with global engineering teams.
- Exposure to impactful cloud reliability projects within secure and regulated environments.