Grafana and Observability Engineer
Dexian · Coppell, TX · 1 wk ago
FinanceFull-time
About the role
We are seeking an experienced Senior Observability Engineer to join our Observability Engineering team. This role focuses on designing, implementing, managing, and automating enterprise-level observability solutions with an emphasis on Grafana, OpenTelemetry, monitoring, alerting, and telemetry operations. The ideal candidate will have deep experience building and running observability platforms at scale, automating infrastructure with Terraform, and helping application and infrastructure teams adopt consistent observability standards. This position will play a major role in modernizing our observability stack and leading the transition from legacy monitoring tools to Grafana-based solutions.
Responsibilities
- Observability Platform Engineering
- Administer and support Grafana Cloud and on-prem Grafana environments.
- Build and deploy observability solutions for metrics, logs, traces, synthetic monitoring, and alerting.
- Define and maintain observability standards, best practices, and governance.
- Configure and manage Grafana data sources, alerting, RBAC, folders, teams, and integrations.
- Ensure the platform is scalable, reliable, resilient, and operationally sound.
- Automation & Infrastructure as Code
- Develop and maintain Terraform modules for Grafana infrastructure and configuration.
- Automate onboarding for applications, dashboards, alerts, and data sources.
- Build self-service capabilities to reduce manual work and improve adoption.
- Integrate observability into CI/CD and infrastructure provisioning workflows.
- Monitoring, Alerting & Incident Management
- Design monitoring and alerting strategies aligned to service health and critical workflows.
- Improve alert quality by reducing noise and strengthening signal accuracy.
- Support incident response, troubleshooting, RCA, and post-incident reviews.
- Continuously enhance operational visibility and platform health.
- OpenTelemetry & Telemetry Engineering
- Implement and support OpenTelemetry instrumentation across applications and infrastructure.
- Define standards for logs, metrics, traces, and telemetry collection.
- Support telemetry pipelines, agent deployments, and data collection strategies.
- Guide teams on instrumentation design and observability adoption.
- Migration & Modernization
- Lead migration efforts from legacy monitoring tools to Grafana.
- Assess existing monitoring, logging, alerting, and tracing setups and recommend modernization paths.
- Build reusable migration patterns, automation, and engineering standards.
- Partner with application teams to accelerate enterprise observability adoption.
- Collaboration & Leadership
- Work closely with development, infrastructure, cloud, and SRE teams.
- Provide technical leadership and mentorship across engineering groups.
- Contribute to observability architecture, strategy, and roadmap.
- Champion observability as a core engineering discipline across the organization.
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
- 5+ years in observability, monitoring, operations, or platform engineering.
- Hands-on experience administering Grafana in large enterprise environments.
- Strong Terraform and Infrastructure as Code (IaC) experience.
- Experience with monitoring, alerting, logging, and distributed tracing.
- Knowledge of OpenTelemetry instrumentation and telemetry pipelines.
- Strong Linux and cloud administration skills.
- Scripting experience with Python, PowerShell, Bash, or similar.
- Understanding of operational excellence, reliability engineering, and incident management.
Preferred Qualifications
- Experience migrating from Splunk, Dynatrace, AppDynamics, New Relic, OpenText OBM, or similar tools.
- Experience with Grafana Alloy, Tempo, Loki, Mimir, or Prometheus.
- Experience running observability platforms in AWS.
- Knowledge of Kubernetes, containers, and cloud-native observability.
- Experience designing enterprise observability strategies and governance.
- Familiarity with CI/CD and DevOps practices.
Skills
- Grafana Administration
- Terraform
- OpenTelemetry (OTEL)
- Monitoring & Alerting
- Observability Engineering
- Platform Engineering
- Linux Administration
- AWS Cloud Services
- Automation & Scripting
- Incident Management
- Infrastructure as Code
- Telemetry Pipelines
- Reliability Engineering
- Root Cause Analysis
- Enterprise Monitoring Architecture