Jobs · Engineering

Senior / Staff Software Engineer, Cloud & Real Time Infrastructure

Antora Energy · United States · 1 wk ago
RemoteRemoteEngineering$183/hrFull-time

About the Role

Antora Energy is seeking an engineer to own the real-time cloud infrastructure connecting our software systems to physical hardware. Every five minutes, our software determines how our thermal batteries should charge and discharge. Those decisions must travel reliably through our messaging infrastructure and reach physical assets in the field. In the other direction, plant telemetry flows from the edge into our cloud data warehouse, where it must be complete, timely, and trustworthy enough to support operational decisions and financial trading. You’ll also own the corporate cloud platform that the rest of the company builds on.

While these systems have very different service-level requirements, they share a need for thoughtful architecture, strong operational practices, and reliable tooling. We’re hiring one person for a scope that many companies divide across multiple teams. This is a high-impact opportunity to establish platform-wide standards, unify observability, and shape our infrastructure and operational strategy from end to end.

Responsibilities

  • Own the Real-Time Telemetry and Dispatch Path: Own the full data path from the edge, through streaming ingestion, into the warehouse, and back out to asset dispatch. You’ll define service-level objectives, build freshness and gap detection, and ensure silent data loss becomes an actionable alert—not something discovered in a report a week later.
  • Manage Our AWS Infrastructure as Code: Own and evolve our AWS footprint across multiple accounts using Terraform. You’ll operate containerized services on ECS and Lambda and manage orchestration, networking, secrets, access controls, and supporting infrastructure. You’ll design systems with failure modes, blast radius, observability, capacity, and recovery in mind from the beginning.
  • Advance an AI-First Software Development Lifecycle: Build the systems that make agent-authored changes safe at scale, including fast and trustworthy CI, hermetic test infrastructure, preview environments, progressive delivery, and automated guardrails. You’ll also establish the cost controls and verification workflows that allow autonomous development systems to operate without requiring a person to supervise every step.
  • Own Reliability for Systems Connected to Physical Assets: Participate in the on-call rotation for systems that dispatch instructions to a physical plant. You’ll build the observability, alerting, and response processes needed to identify problems before they affect field operations. You’ll also define clear escalation paths across cloud infrastructure, site controls, and our market operations desk.
  • Build a Platform Engineers Want to Use: Treat infrastructure as an internal product. Create self-service tools, paved paths, and platform capabilities that help engineers ship safely without unnecessary friction. You’ll set standards where consistency matters while avoiding process for process’s sake.

Requirements

  • Deep production infrastructure experience: Several years of experience building and operating cloud infrastructure at scale, with strong expertise in AWS, Terraform, containers, and CI/CD.
  • Hands-on site reliability experience: Carried primary on-call responsibility for a production system that mattered. Planned and completed at least one zero-downtime migration involving a stateful, always-on service.
  • Experience with real-time or physical systems: Worked in industrial technology, energy, robotics, automotive, or another environment where software failures have tangible operational consequences.
  • An internal-product mindset: Built self-service tooling that engineers chose to adopt—not simply tooling they were required to use. Understand how to balance standardization, usability, and developer autonomy.
  • Strong programming skills: Strong coding skills in Python, Go, or both. Do more than provision infrastructure: write and operate the software that keeps systems reliable.
  • Practical experience with AI development tools: Use AI coding tools regularly and have a thoughtful perspective on where they work well, where they fail, and how their output should be reviewed and verified.
  • Sound operational judgment: Can distinguish between a safeguard that prevents a meaningful failure and a process that only creates friction. Design controls proportionate to the risks involved.
  • A willingness to learn across the stack: Deep experience with industrial protocols is valuable but not required. A strong systems engineer who is excited to learn the plant and operational technology side of the business can thrive in this role.

Skills

  • Experience with industrial and operational technology protocols, including MQTT or Sparkplug B
  • Familiarity with platforms such as HiveMQ, NATS, Ignition, or industrial historians
  • Experience operating across strict OT/IT boundaries
  • Experience with air-gapped or intermittently connected environments
  • Expertise in time-series or high-cardinality data at scale
  • Experience with ClickHouse, Dagster, or streaming ingestion systems
  • Experience building developer tooling specifically for AI coding agents
  • Familiarity with sandboxing, AI cost governance, or agent-readable test output
  • Experience with energy markets, dispatch systems, or grid-connected assets

Pay

Salary Range: $183,000–$240,000 USD. The actual salary offered will be determined based on a candidate's experiences, credentials, and expertise.

Benefits

  • Equity compensation in the form of stock options
  • Premium health benefits package with life and disability insurance
  • 401K plan with employer contributions
  • Flexible spending accounts
  • Industry-leading paid-time-off policy featuring flexible and inclusive holiday observance
  • Paid volunteer time off

Schedule

Work Location: Remote / US based

Similar jobs