Senior Observability Operations Engineer
Santcore Technologies · Phoenix, AZ · 1 mo ago
On-siteEngineeringContract
Key Responsibilities
- Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
- Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
- Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
- Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
- Develop dashboards, alerts, reports, and executive operational metrics.
- Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
- Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
- Perform root cause analysis for production incidents using observability platforms.
- Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
- Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.
- Participate in incident, problem, change, and release management processes.
- Drive platform upgrades, patching, security compliance, and operational governance.
- Improve platform reliability through automation, self-healing, and AI-assisted operations.
Required Technical Skills
- Observability Platforms
- Dynatrace Administration
- Splunk Enterprise Administration
- OpenSearch Administration
- Elasticsearch Administration
- Grafana
- Prometheus
- Kibana
- Jaeger
- OpenTelemetry
- Kafka (preferred)
- Infrastructure
- Linux Administration
- Kubernetes
- Docker
- OpenShift or Rancher
- Networking (TCP/IP, DNS, Load Balancers, Firewalls)
- System Administration
- Cloud & DevOps
- AWS, Azure, or GC
- PCI/CD pipelines
- Git
- Terraform
- Ansible
- REST APIs
- Scripting
- Python
- Bash/Shell
- PowerShell (preferred)
- AI & Automation Skills
- Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency.
- Knowledge of AIOps platforms and AI-driven observability.
- Experience with Dynatrace Davis AI for anomaly detection and root cause analysis.
- Understanding of machine learning concepts for predictive monitoring and intelligent alerting.
- Experience building AI-assisted operational runbooks and troubleshooting workflows.
- Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and AI powered knowledge search.
- Familiarity with LLM's, prompt engineering, and AI assisted automation.
- Experience using Python with AI frameworks (LangChain, LangGraph, Open AI APIs or similar).
- Exposure to AI driven incident summarization, log analysis, and automated ticket enrichment.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering or equivalent experience.
- 6-10+ years of IT infrastructure or observability operations experience.
- 4+ years administering Dynatrace, Splunk, OpenSearch or ElasticSearch.
- Strong Linux system admin experience.
- Experience supporting enterprise scale production environments.
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management skills.