Sr. Software Engineer
Fullbay · Phoenix, AZ · 1 wk ago
EngineeringFull-time
Primary Duties & Responsibilities
- Design, build, and maintain observability infrastructure including distributed tracing, structured log aggregation, metrics pipelines, and alerting systems across Fullbay Next microservices using CloudWatch, OpenTelemetry, and AWS-native services
- Build and operate AI-assisted anomaly detection capabilities that help engineering teams identify and resolve production issues faster, reducing mean time to detection and resolution across the platform
- Develop internal developer tooling and self-service operational dashboards that surface real-time system health, cost visibility, and reliability signals, empowering engineers to own their services in production
- Establish and govern Terraform standards for provisioning observability infrastructure, writing reusable modules and reviewing infrastructure-as-code contributions across engineering teams
- Define and enforce instrumentation standards so that every service on Fullbay Next emits consistent, high-quality telemetry from day one, making log aggregation, tracing, and alerting reliable and low-maintenance across the platform
- Drive FinOps visibility by building cost and usage dashboards that give engineering and leadership clear, actionable insight into AWS spend at the service level
- Define and enforce SLI, SLO, and error budget frameworks and build the systems that measure and report on them, enabling teams to make data-driven reliability tradeoffs
- Collaborate with Sherpa platform team members to integrate observability tooling into the internal developer platform, making provisioning, deployment, and operational monitoring seamless for every engineering team
- Investigate and evaluate emerging observability technologies and AI tooling, making informed build-vs-buy recommendations that align with Fullbay's AWS-first, serverless-preferred architecture
- Adhere to all confidentiality and compliance regulations
Minimum Education & Work Experience
- This job requires at least 7-10 years of experience in Software Design and Development, with significant depth in observability engineering, platform engineering, or developer support operations in a cloud-native environment.
Key Skills And Qualifications
- Hands-on experience with CloudWatch and OpenTelemetry for distributed tracing, structured logging, metrics collection, and alerting across microservices architectures
- Strong proficiency in Java, TypeScript/Node.js, and Python; experience building production services on AWS Lambda, ECS, and DynamoDB in a 100% AWS environment
- Terraform expertise: writing, maintaining, and governing infrastructure-as-code for observability systems and developer platform components
- Experience building AI-assisted operational tooling, including anomaly detection, log analysis, or intelligent alerting systems that reduce manual triage burden
- Deep familiarity with FinOps principles and AWS cost visibility tooling, with the ability to build service-level cost dashboards that drive accountability across engineering teams
- Awareness of SRE best practices including SLI/SLO definition, error budgets, and reliability-first service design within a microservices platform
- Strong ability to define and communicate instrumentation standards, developer platform conventions, and operational readiness requirements across engineering teams of varying experience levels