Jobs · Engineering · Georgia

Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert

RaceTrac · Atlanta, GA · 2 wk ago
EngineeringFull-time

Dynatrace & Observability Engineering

Serve as the primary Dynatrace SME across the organization.
Design, develop, and optimize enterprise observability solutions using Dynatrace.
Develop advanced Dynatrace DQL queries, dashboards, workflows, alerts, and analytics.
Implement intelligent monitoring strategies for applications, APIs, integrations, Azure services, mobile platforms, and distributed systems.
Continuously improve observability maturity through telemetry standardization, proactive monitoring, and automation.
Configure and tune alerting mechanisms to improve signal-to-noise ratio and reduce alert fatigue.
Enable and enhance observability for mobile applications across iOS and Android platforms.

Azure Monitoring & Cloud Operations

Build and maintain monitoring solutions using:
Azure Monitor
Application Insights
Azure Log Analytics
Azure KQL
Monitor and troubleshoot Azure Function Apps, App Services, APIs, integrations, and backend services.
Analyze telemetry, traces, logs, metrics, and distributed transactions to identify root causes and performance bottlenecks.
Troubleshoot cloud-native applications and Azure infrastructure issues.
Develop proactive monitoring for cloud services, integrations, APIs, and backend processing systems.

API & Integration Monitoring

Monitor and troubleshoot Azure API Management (APIM), API Gateways, API endpoints, and integrations.
Understand end-to-end API transaction flows and dependency mapping.
Build observability solutions for APIs, middleware platforms, and integration services.
Diagnose latency issues, transaction failures, authentication issues, and backend service degradation.

Mobile Application Observability

Enable telemetry, monitoring, tracing, and performance analysis for iOS and Android applications.
Analyze mobile-to-backend transaction flows and end-user experience metrics.
Troubleshoot mobile application latency, crash analytics, API failures, and connectivity issues.
Correlate mobile telemetry with backend application and infrastructure monitoring data.

Application Engineering & Troubleshooting

Utilize prior .NET development experience to troubleshoot application behavior, performance, and deployment issues.
Read and understand .NET application code to support root cause analysis and observability implementation.
Work closely with development teams to understand application logic, API flows, dependencies, and exception handling.
Support Azure Function deployments, configuration management, scaling, and runtime troubleshooting.
Collaborate with development teams during architecture reviews and production releases.
Ensure observability and monitoring readiness before deployments go live.

Site Reliability Engineering (SRE)

Perform deep technical analysis across systems by correlating logs, metrics, traces, and application telemetry.
Conduct root cause analysis (RCA) for recurring incidents and systemic issues.
Partner with engineering and operations teams to implement preventive improvements and automation.
Develop KPI-driven reliability improvements focused on system stability, performance, and operational excellence.
Proactively identify risks, bottlenecks, failure patterns, and reliability concerns before business impact occurs.

Continuous Improvement & Automation

Automate operational workflows and monitoring processes wherever possible.
Improve operational efficiency using AI-driven insights and automation capabilities.
Build reusable monitoring frameworks, dashboards, and telemetry standards.
Drive observability best practices across engineering teams.

Similar jobs