Splunk Observability Engineer
Diligent Tec, Inc · Chicago, IL · 1 wk ago
On-siteInformation TechnologyContract
Location: Chicago, IL (onsite)
About the role
Design, implement, and optimize a full-stack observability strategy using the Splunk Observability Cloud (formerly SignalFx) and Splunk Enterprise/Cloud. Ensure that engineering teams have 360-degree visibility into system health, moving the organization from reactive "firefighting" to proactive "pattern-based" incident prevention.
Responsibilities
- Architect the ingestion of the "Three Pillars" (Metrics, Logs, Traces) using OpenTelemetry (OTel) collectors.
- Develop logic to aggregate high-cardinality data to reduce "noise" while maintaining "signal" for troubleshooting.
- Use SPL (Search Processing Language) and SignalFlow to perform pattern analysis, detecting anomalies before they trigger traditional threshold alerts.
- Build executive and technical dashboards that correlate disparate data points (e.g., showing how a spike in 500-errors in Logs relates to a specific span in a Trace).
Requirements
Telemetry & Data Specialization
- Logs: Proficiency in "Logging-in-Context." Ability to link logs directly to trace IDs so developers can jump from a failing trace to the specific line of code in the logs.
- Metrics: Expertise in SignalFlow (Splunk’s background streaming analytics language). Knowledge of calculating percentiles (P95, P99), rates of change, and historical averages.
- Traces: Deep understanding of Distributed Tracing. Ability to instrument applications (Java, Python, Go) to capture spans and identify bottlenecks in microservices.
Pattern Analysis & Aggregation
- Anomaly Detection: Ability to configure Metric Finder and MDetector using standard deviations or "Mean Absolute Deviation" to find outliers.
- Data Scrubbing: Skills in using Splunk Ingest Actions or Edge Processors to filter, mask, or aggregate data at the edge to save on license costs and improve search speed.
- Pattern Discovery: Using Splunk’s machine learning commands (e.g., findkeywords, cluster) to group millions of log events into a few dozen "patterns" for faster root cause analysis.
Dashboards & Visualization
- High-Cardinality Handling: Designing dashboards that don’t "break" when viewing thousands of containers.
- Contextual Drill-downs: Building "Glass Tables" (in ITSI) or Unified Dashboards that allow a user to click a metric and immediately see the associated logs.
- Frameworks: Familiarity with the Dashboard Studio and JSON-based dashboard definitions for version control (GitOps).
Qualifications
- DevOps & IAC skills
- Splunk Cloud Certified Metrics User (focuses on the metrics and alerting side)
- Splunk Core Certified Power User (essential for mastering complex SPL for log analysis)
- OpenTelemetry Expert: Knowledge of the OTel Collector configuration (receivers, processors, exporters) is highly desirable.