Software Engineer, Infrastructure
The Role
Own and evolve our Kubernetes infrastructure, including cluster management, service mesh configuration, and container security policies.
Design and implement progressive delivery pipelines with canary deployments, automated rollbacks, and deployment health validation.
Build and maintain our observability infrastructure in Datadog, including dashboards, monitors, SLOs, and distributed tracing.
Drive incident response for high-severity outages and proactively model capacity needs for low-latency AI inference.
Architect and automate secure infrastructure using Infrastructure-as-Code for VPCs, IAM policies, Kubernetes manifests, and private cloud deployments.
Maintain and improve the infrastructure controls that support our SOC 2 compliance posture.
Lead customer engagements for enterprise rollouts and mentor mid-level engineers on infrastructure best practices.
Must-Haves
- 8+ years in infrastructure engineering or DevOps at high-growth or hyperscale companies.
- Experience with Docker and Kubernetes, including production cluster management, Helm, and service mesh technologies.
- A proven track record of architecting and operating AWS (preferred), GCP, or Azure at an enterprise scale.
- Experience with observability platforms, preferably Datadog (metrics, logs, APM, distributed tracing).
- A strong background in Infrastructure-as-Code (Terraform, Helm, Kustomize) and safe deployment practices (progressive delivery, canary deployments, GitOps, automated rollbacks).
- "Battle scars" from leading outages, capacity events, and large-scale incident reviews.
- Strong programming skills in Python.
Bonus Points
- Familiarity with TypeScript.
- Direct involvement in SOC 2 or other compliance audit preparation or remediation.
- Previous experience at startups scaling infrastructure from the early stages to the enterprise level.
- Direct experience with private-cloud or on-premises deployments for regulated customers.
- A background in fintech or building systems for highly regulated industries.
- Experience with AI/ML infrastructure and model deployment at scale.