Staff Software Engineer - AIOps
About the role
The Staff Software Engineer - AIOps is a senior hands-on technical contributor responsible for designing, building, and operating the AIOps agent platform that enables AI agents to observe events across Early Warning technology environments, generate diagnoses and recommendations, and execute approved actions within defined safety controls.
Responsibilities
- Apply mature software engineering practices across the agent runtime, tool layer, APIs, and operator experience, including versioned interfaces, automated testing, code review, release management, observability, and regression coverage.
- Execute complex AIOps engineering assignments with general direction; break work into deliverable components, identify practical implementation solutions, raise risks and dependencies, and seek guidance when decisions extend beyond established architecture or standards.
- Contribute to the design and evolution of the agent platform architecture; maintain architecture decision records, define interface standards, and evaluate significant design and build-versus-buy tradeoffs with appropriate technical guidance.
- Design, build, test, and operate the core agent runtime, including event ingestion, context assembly, planning, tool invocation, verification, escalation, and response handling.
- Design and maintain scoped, typed, and auditable tool interfaces that enable agents to interact with CI/CD platforms, infrastructure automation, Kubernetes, network services, ITSM platforms, observability systems, and secrets management solutions.
- Define and implement controls for agent actions, including read-only and advisory capabilities, human approval requirements, narrowly scoped unattended actions, permission boundaries, blast-radius limits, dry-run capabilities, reversibility, and emergency shutdown controls.
- Define tool contracts, versioning standards, testing requirements, and intended behavior so agent-accessible capabilities are managed as reliable production interfaces.
- Establish telemetry capabilities used by the agent, including logs, metrics, and traces from relevant technology platforms, and contribute to standards for telemetry quality and schema evolution.
- Design and operate model routing, session and state management, retries, timeouts, and cost and latency controls appropriate for a production service.
- Implement comprehensive auditability for events received, decisions generated, approvals obtained, and actions executed to support operational, risk, compliance, and examination requirements.
- Develop evaluation capabilities for agent decision quality, including historical incident replay, controlled or shadow-mode evaluation, measurement of proposed actions, and regression testing.
- Develop feedback mechanisms that incorporate human approvals, overrides, and operational outcomes into measurable improvements to platform quality.
- Package reusable agent capabilities, tool integrations, APIs, documentation, and implementation patterns so other technology teams can extend and adopt the platform through defined self-service practices.
- Develop and maintain the operator experience, including approval workflows, agent activity and decision history, telemetry views, and APIs that expose agent state to user interfaces and other systems.
- Partner with Engineering, Technology Operations, Architecture, Security, Risk, and Compliance stakeholders to evaluate technical tradeoffs and ensure solutions meet reliability, security, operational, and regulatory requirements.
- Define and monitor measures of platform effectiveness, including recommendation quality, approval and override rates, action success, rollback frequency, latency, cost, adoption, and operational outcomes.
Requirements
Typically eight or more years of related experience in software engineering, platform engineering, distributed systems, DevOps, AIOps, or a related technical discipline.
Demonstrated strength in software engineering and distributed-systems design, including the ability to turn ambiguous technical problems into maintainable, testable solutions and reason about interfaces, state, retries, timeouts, idempotency, failure modes, and operational risk.
Demonstrated ability to execute complex technical work with general direction, independently make implementation decisions within established architecture and standards, and recognize when guidance or escalation is needed for novel or enterprise-wide decisions.
Strong production development experience with Python, including services, APIs, automation, orchestration systems, or agent runtimes.
Hands-on experience developing production agentic systems, tool-using AI applications, or autonomous or semi-autonomous software control loops beyond proof-of-concept implementations.
Experience designing tool-use or function-calling interfaces for AI agents using MCP or comparable integration patterns.
Experience implementing production controls for AI or automated systems, including evaluation, guardrails, monitoring, cost and latency management, and management of unintended or failed actions.
Experience with event-driven architectures and messaging technologies such as Kafka, Pub/Sub, RabbitMQ, or comparable platforms.
Experience with production observability practices, including structured logging, metrics, distributed tracing, and technologies such as OpenTelemetry, Prometheus, or Grafana.
Experience developing front-end solutions using modern frameworks such as React, Next.js, or Angular, along with backend or API development using TypeScript or Python.
Demonstrated ability to work effectively in an environment requiring strong controls, explainability, approval processes, and auditability for automated actions.
Demonstrated ability to lead complex technical work, collaborate across engineering and operations disciplines, and influence implementation decisions without direct authority.
Strong written and verbal communication skills, including technical documentation, architecture decision records, implementation standards, and control documentation.
Current AWS, Kubernetes, AI engineering, security, or other relevant technical certification.
Qualifications
Education and/or experience typically obtained through a bachelor's degree in computer science, engineering, or a related technical field.
Ability to independently possess the eligibility to work in the United States at the date of hire.
Eligible for a discretionary incentive plan and benefits.