Open Telemetry (OTel) Instrumentation Engineer
VOLTO Consulting · Los Angeles, California, United States · Yesterday
On-siteEngineeringContract
Role: Senior Open Telemetry (OTel) Instrumentation Engineer Descriptions "Core Responsibilities SaaS Instrumentation End-to-End Instrumentation: Instrument SaaS products utilizing OTel SDKs (Python, Node.js, Java, Go) to ensure spans, metrics, and structured logs are emitted at the appropriate granularity. Semantic Conventions: Define and enforce attribute naming conventions (e.g., service.name, tenant.id, feature.flag) aligned strictly with OTel semantic standards. Multi-Tenant Observability: Instrument multi-tenant surfaces to guarantee robust tenant-level observability while preventing cross-tenant data leakage. Cloud Infrastructure Collector Fleet Management: Deploy and maintain the OTel Collector fleet across AWS, GCP, and Azure, including receiver configurations, processor pipelines, and exporter routing. Runtime Instrumentation: Instrument serverless (Lambda, Cloud Run) and container runtimes (EKS, GKE, AKS) utilizing auto-instrumentation where feasible, and manual instrumentation when technical nuance requires it. Metric Normalization: Collect and normalize cloud provider metrics (CloudWatch, Cloud Monitoring, Azure Monitor) via OTel receiver plugins. Network Telemetry Data Collection: Gather network flow data (sFlow, NetFlow/IPFIX), SNMP traps, and BGP state via OTel-native and bridged receivers. Context Propagation: Correlate network events with application traces using consistent trace-context propagation across network boundaries. Strategy Collaboration: Partner with NetOps to define the MELT (Metrics, Events, Logs, Traces) strategy for on-prem, SD-WAN, and cloud interconnect segments. Application Observability APM Ownership: Own Application Performance Monitoring (APM) instrumentation, including distributed tracing, custom span attributes, database query capture, and error fingerprinting. Asynchronous Workloads: Instrument async workloads (Kafka, SQS, Celery) ensuring W3C Trace Context propagation across message boundaries. SLI/SLO Definition: Define and instrument Service Level Indicators (SLIs) and Service Level Objectives (SLOs) such as latency histograms, error-rate counters, and availability gauges queryable directly from the Lakehouse. Platform & Standards Internal Libraries: Build and maintain internal instrumentation libraries and language-specific wrappers that encode organizational conventions. Governance: Author and review instrumentation RFCs while driving the adoption of OTel semantic conventions across all engineering teams. Data Alignment: Collaborate with the AIOps ML team to ensure telemetry schemas meet feature engineering and ML model requirements. Pipeline Operations: Operate the Collector pipeline at scale, managing backpressure handling, sampling strategies (tail/head), and cardinality budgets." "Strong expertise in: OpenTelemetry (SDKs, APIs, Collector, OTLP) Distributed tracing, metrics, logging principles [opentelemetry.io]