Lead Performance & Observability Engineer
About the role
Intercontinental Exchange, Inc. (ICE), the owner of the New York Stock Exchange (NYSE), is seeking a results-oriented, self-motivated individual for its Clearing Performance Engineering team in Atlanta. This role serves as a technical lead within a team of software architects and performance engineers, operating in a cutting-edge technology environment responsible for running critical financial sector exchanges and clearinghouses.
The successful candidate will drive performance engineering strategy across multiple platforms, mentor peers, and deliver deep-dive analysis on the most complex and time-sensitive systems in the organization. This role requires close collaboration with software architects, developers, quant library owners, infrastructure teams, and project managers to ensure platforms meet the highest standards of reliability, scalability, and throughput in a mission-critical environment where end-of-day processing windows are measured in minutes.
Responsibilities
- Strategy & Investigation
- Define and own the performance engineering strategy across multiple critical platforms; set standards for testing approach, tooling, and reporting.
- Lead deep-dive performance investigations on CPU-bound, multi-threaded Java systems—including analysis at the JVM, OS, and hardware levels.
- Drive scalability analysis: model system performance as data volume and instrument counts grow; produce capacity projections to guide architecture decisions.
- Lead critical path segregation analysis: identify time-constrained operations, propose architectural solutions (e.g., head-start strategies, out-of-band processing), and validate their impact.
- Observability & Monitoring Framework
- Design and own a comprehensive observability framework spanning metrics, distributed tracing, and continuous profiling for full-stack visibility.
- Build and maintain platform-wide health scoring and alerting using Prometheus and VictoriaMetrics; drive consistent instrumentation and Grafana visualization standards.
- Develop custom metrics exporters to surface business-layer signals (database growth, queue depths, pipeline latency) alongside infrastructure metrics.
- Architect and maintain Grafana dashboard ecosystems; establish standards for dashboard organization, datasource governance, and operational visibility.
- Lead distributed tracing and continuous profiling initiatives; use trace and profile data to complement load test findings and accelerate root cause analysis.
- Build automated post-test analysis pipelines correlating load test results with observability data to produce clear performance narratives for stakeholders.
- Testing & Automation
- Design and build robust test harnesses to measure version-to-version performance regressions for compute-intensive components.
- Tune JVM thread pools, garbage collection, and heap allocation for high-throughput, latency-sensitive processing pipelines.
- Design workloads representative of real production traffic patterns using load generation frameworks (JMeter, Gatling, or custom harnesses).
- Create and maintain automation scripts and tooling to simplify repeatable performance analysis tasks.
- Collaboration & Leadership
- Act as a technical liaison between performance, development, infrastructure, and operations teams; translate findings into actionable recommendations with clear data backing.
- Champion AI-augmented workflows: identify where AI tools accelerate root cause analysis, reduce boilerplate in test harnesses, and surface anomalies in benchmark data while establishing validation standards.
Requirements
- Bachelor's Degree or equivalent in Computer Science, Engineering, or a related field.
- 8+ years of experience in performance engineering, performance testing, or Java development in high-volume, low-latency, transactional systems.
Skills
- JVM & Application Performance
- Deep expertise in JVM internals: heap dump analysis, thread dump analysis, GC log interpretation, and memory profiling.
- Proficiency with Java and scripting languages (Python, Groovy, or Linux shell) for building test harnesses and automation.
- Proven ability to perform scalability analysis—projecting system behavior as data volumes grow and validating linearity assumptions.
- Strong grasp of critical path analysis: identifying time-constrained operations and validating architectural changes to improve throughput.
- Observability Stack & Monitoring Frameworks
- Hands-on experience designing multi-signal observability stacks (metrics, distributed tracing, continuous profiling) for Java applications.
- Proficiency with Prometheus and VictoriaMetrics for metrics collection, storage, querying, health scoring, and alerting.
- Experience building custom metrics exporters to expose business-layer and application-specific signals.
- Proficiency with Grafana for operational and performance dashboards; experience with programmatic dashboard management and datasource governance.
- Familiarity with distributed tracing and continuous profiling tools; ability to use trace and profile data alongside load test results.
- Experience building automated post-test reporting pipelines synthesizing observability data and load test results.
- Messaging, Load Testing & Event Pipelines
- Experience with event-driven, message-based architectures (Kafka, IBM MQ, or equivalent); ability to test and profile throughput and latency across event pipelines.
- Experience with load generation frameworks (JMeter, Gatling, or custom harnesses) and designing representative workloads.
- Communication & AI Tooling
- Excellent verbal and written communication skills; ability to present findings clearly to technical and non-technical stakeholders.
- Ability to collaborate across teams with varying levels of performance domain knowledge.
- Demonstrated proficiency with AI coding and analysis tools (GitHub Copilot, Claude, Cursor, or equivalent) for performance-specific tasks (profiling scripts, Prometheus queries, GC log analysis, test harness scaffolding).