Software Developer 5
Oracle · Nashville, TN · 3 wk ago
Engineering$135k–$306k/yrFull-time
Responsibilities
- Lead the design, development, and operation of cloud-scale observability platforms supporting metrics, logs, traces, and related telemetry data.
- Architect and implement highly scalable, resilient, and cost-efficient telemetry collection, ingestion, processing, storage, and query systems.
- Drive the evolution of end-to-end observability pipelines, from instrumentation and data collection through real-time analytics and long-term retention.
- Design and optimize distributed systems capable of ingesting and processing massive volumes of telemetry data with stringent latency and availability requirements.
- Develop scalable storage and indexing solutions for high-cardinality metrics, large-scale log analytics, and distributed tracing workloads.
- Build and enhance query, search, and retrieval services that deliver fast, reliable, and intuitive access to observability data.
- Collaborate with product management, architects, SREs, and engineering teams to define and deliver next-generation observability capabilities.
- Identify and resolve performance bottlenecks across the observability stack, including ingestion, storage, indexing, aggregation, and query execution.
- Drive technical strategy and architectural decisions for observability services operating at hyperscale cloud environments.
- Mentor senior and junior engineers, provide technical leadership, and foster engineering best practices across the organization.
- Partner with service teams to improve instrumentation, telemetry quality, and operational visibility across cloud services.
- Establish and monitor key service health, scalability, performance, and cost-efficiency metrics for observability platforms.
- Lead troubleshooting and root-cause analysis efforts for complex distributed systems and large-scale production environments.
- Stay current with emerging trends, technologies, and best practices in observability, distributed systems, data processing, and cloud-native architectures.
Qualifications
- Passionate about building cloud-native observability platforms that power both the cloud itself and the customers who depend on it.
- Experience in designing, developing, and operating cloud-scale observability platforms.
- Strong background in distributed systems, telemetry collection, processing, storage, and query systems.
- Knowledge of high-throughput telemetry ingestion, large-scale data processing, cost-efficient storage, low-latency query execution, multi-tenant reliability, and operational excellence.
- Ability to develop scalable storage and indexing solutions for high-cardinality metrics, large-scale log analytics, and distributed tracing workloads.
- Experience in building and enhancing query, search, and retrieval services for observability data.
- Collaborative skills with product management, architects, SREs, and engineering teams.
- Experience in identifying and resolving performance bottlenecks across the observability stack.
- Experience in driving technical strategy and architectural decisions for observability services operating at hyperscale cloud environments.
- Ability to mentor senior and junior engineers, provide technical leadership, and foster engineering best practices.
- Experience in partnering with service teams to improve instrumentation, telemetry quality, and operational visibility across cloud services.
- Experience in establishing and monitoring key service health, scalability, performance, and cost-efficiency metrics for observability platforms.
- Experience in leading troubleshooting and root-cause analysis efforts for complex distributed systems and large-scale production environments.
- Ability to stay current with emerging trends, technologies, and best practices in observability, distributed systems, data processing, and cloud-native architectures.