Jobs · Engineering · Washington

Senior Platform Operations Engineer, Infrastucture

Viome · Bellevue, WA · Yesterday
On-siteEngineeringFull-time

About the role

We are looking for a Senior Platform Operations Engineer, Infrastructure to take ownership of Viome’s Azure-based service platform with a clear mandate: make it dramatically simpler through robust Unix-based systems engineering. This is a consolidation role for a seasoned systems administrator focused on migrating a suite of abstracted services onto a deliberately stable target platform: Linux VMs, systemd-supervised services, Apache, HAProxy, and Nginx routing. You must possess strong Unix knowledge to operate and eventually transform the current stack safely, with a bias toward simplicity and host-level stability.

Responsibilities

  • Simplification & Migration (core mandate)
    • Plan and execute zero-downtime migrations of services from AKS/containers to VM-based hosting: systemd unit authoring and service supervision (restart policy, resource limits, sandboxing), Apache, HAProxy and/or nginx as reverse proxy and TLS terminator, certificate automation.
    • Build and own the deployment scripts for the target platform: build → test → archive → ship → symlink-flip → health-check → rollback, scripted in bash.
    • Preserve two non-negotiable invariants while simplifying everything else: immutable, commit-traceable artifacts and scripted rollback.
    • Inventory and retire stale infrastructure: dormant deployments, unused DNS records, orphaned firewall rules, unpinned image tags.
    • Maintain host hygiene for consolidated services: OS patching discipline, runtime vendoring, log rotation, centralized log aggregation. Familiarity with UptimeKuma, Nagios or similar.
    • Operate and migrate data stores: PostgreSQL and/or MySQL.
  • Network and Infrastructure Operations
    • Operate and troubleshoot workloads during the transition, focusing on the networking layer, load balancing, and core platform services.
    • Administer the hub network: Azure Firewall rules, VPN gateways, VNet peering, public and private DNS zones and reason about a packet’s full path from public IP to service.
    • Support the existing release process and network and infrastructure operations until each service is migrated to the new VM-based standard.
    • Keep the observability stack healthy (OpenTelemetry, ELK, Grafana, uptime and cost monitoring) and carry its essentials forward to the simplified platform.
    • Coordinate cross-cloud dependencies with AWS: DNS/edge routing, queue consumers, and egress IP allowlists.
    • L2+ operations support; manage runbooks for external L0, L1 support.
  • External Integrations & Security
    • Own the external integrations most at risk during migration: e-commerce and subscription platforms, messaging/notification providers, and clinical/health-data partners — webhook delivery, signature verification, idempotency, and retry semantics.
    • Raise the security baseline as you consolidate: secrets management, webhook authentication, least-privilege network access, and data-retention hygiene.

Requirements

  • Required
    • 7+ years operating production Unix/Linux systems, with deep knowledge of systemd, process supervision; Apache, HAProxy.
    • Strong shell plus one scripting language (bash/Python/PHP) with a track record of building deploy and rollback tooling, not just using it.
    • Demonstrated reverse-engineering ability: taking ownership of an undocumented production service and recovering its real dependencies and failure modes.
    • Message-queue literacy: Azure Service Bus, SQS, RabbitMQ, or equivalent — ordering, lock/ack semantics, dead-letter handling.
    • PostgreSQL operations: replicas/clones, connection proxying, network-restricted access, production diagnostics.
    • Working proficiency with Kubernetes and cloud networking — enough to operate private AKS clusters, an Istio-style ingress layer, and Azure firewall/DNS during the transition.
    • Migration experience: consolidating or re-platforming production services with zero-downtime cutover and tested rollback.
    • A demonstrable record of reducing system surface area — services consolidated, infrastructure retired.
    • Clear written communication for runbooks, migration plans, and cross-functional coordination.
  • Strongly Preferred
    • Azure networking at landing-zone depth: hub-spoke VNets, Azure Firewall, Private Link/Private DNS, VPN gateways.
    • Administration of external-dns, cert-manager, and policy engines such as Kyverno.
    • ELK and OpenTelemetry pipeline operations.
    • Webhook-heavy integration experience (e-commerce platforms such as Shopify, subscription billing, marketing/notification platforms, or healthcare data exchanges).
    • Experience in a regulated or health-data environment: PHI handling, audit trails, least-privilege network design.
    • Terraform or equivalent IaC for cloud network and compute resources.
    • Enough AWS to manage cross-cloud seams (Route53, SQS, egress allowlists).
    • AI minded; proficient with Claude or similar.

What Success Looks Like

  • First 30 days: Trace and fix a production issue end-to-end (DNS → firewall → ingress → service → database) with guidance; produce a true-dependency inventory for one candidate service.
  • First 90 days: First service migrated off AKS to the VM/systemd platform with a passing rollback drill; firewall rules, DNS zones, and cluster add-ons documented and reproducible; stale-resource inventory complete and retirement underway.
  • First 6 months: Migration cadence established with multiple services consolidated; release and rollback drills routine on both platforms; measurable reduction in infrastructure footprint and spend; security-baseline remediations closed.

Similar jobs