Jobs · Colorado

Senior Infrastructure Engineer - AI Ops

Ecommerce Guide · Austin, CO · 1 wk ago
Hybrid$136k–$204k/yrFull-time

At Commerce, our mission is to empower businesses to innovate, grow, and thrive with our open, AI-driven commerce ecosystem. As the parent company of BigCommerce, Feedonomics, and Makeswift, we connect the tools and systems that power growth, enabling businesses to unlock the full potential of their data, deliver seamless and personalized experiences across every channel, and adapt swiftly to an ever-changing market.

We’re hiring a Senior Infrastructure Engineer to join Commerce’s AI Operations team — a highly autonomous team focused on rapid development of AI tools. We operate like a startup within the organization, with a small headcount and full ownership from idea through production. You’ll own the infrastructure, delivery pipelines, and cloud architecture that enable every AI automation.

About the role

This is a hands-on infrastructure role. You’ll manage cloud infrastructure and IaC, build and maintain CI/CD pipelines, configure networking and load balancing, harden security posture, and instrument everything with structured observability. You’ll partner closely with the Data Solutions, Infrastructure and Application Security teams to keep shared infrastructure well-documented.

Responsibilities

  • Design, provision, and manage cloud infrastructure on GCP (Cloud Run, GKE, Redis/Memorystore) using Terraform and Puppet.
  • Build and maintain CI/CD pipelines in Cloud Build, GitHub Actions, and CircleCI—including automated testing gates, canary and blue-green deployments, and artifact promotion across environments.
  • Design and maintain the full edge-to-origin traffic stack: DNS, load balancing (NGINX, GCP Cloud Load Balancing), CDN, WAF, edge compute (Cloudflare, Cloud Armor), bot management, authentication, and rate limiting.
  • Operate and tune container orchestration for AI workloads—autoscaling policies, resource quotas, cold-start optimization, and right-sizing compute for inference cost efficiency.
  • Implement production observability end-to-end: structured logging (Cloud Logging, ELK/Kibana), metrics and dashboards (Prometheus, Grafana), distributed tracing, SLO/SLI definition, and alerting. Own incident response coordination and blameless post-mortems.
  • Implement strong security posture—least-privilege IAM with Workload Identity, network segmentation, firewall rules, and VPC Service Controls.
  • Evaluate, build, and integrate infrastructure tooling—whether that’s custom automation in Python or a managed AI serving platform.

Requirements

  • 7+ years in infrastructure engineering, DevOps, SRE, or platform engineering with direct operational ownership of production systems.
  • Deep Linux systems administration experience (Debian, Ubuntu) at scale.
  • Strong hands-on experience with GCP services (Cloud Run, GKE, Cloud Build, IAM, VPC).
  • Infrastructure as code as a core discipline (Terraform, Puppet).
  • Experience designing and maintaining non-trivial CI/CD delivery pipelines. Cloud Build, GitHub Actions, CircleCI.
  • Hands-on experience with NGINX and edge infrastructure at scale, including DNS, load balancing, and CDN (Cloudflare or equivalent).
  • Hands-on experience with distributed monitoring and logging stacks—Prometheus, Grafana, ELK/Kibana, Cloud Logging.
  • Proficiency in Python, Bash, and shell scripting for infrastructure automation and custom tooling.
  • You design infrastructure with least-privilege IAM, Workload Identity, and compliance controls as defaults.

Nice to have

  • Experience with ML infrastructure.

Pay

The national base salary range for this role is $135,960 - $203,940. Final compensation will be determined based on factors such as relevant experience, skills, qualifications and geographic location. This role may also be eligible for variable compensation (such as bonus or commission), equity, and benefits in accordance with local policies.

Similar jobs