Observability Lead-Sr. Infrastructure Engineer
About the role
The position is described below. If you want to apply, click the Apply Now button at the top or bottom of this page. After you click Apply Now and complete your application, you'll be invited to create a profile, which will let you see your application status and any communications. If you already have a profile with us, you can log in to check status. Need Help? If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).
Responsibilities
- Establish the vision and direction for metrics, logs, traces, telemetry pipelines, event streaming, and developer-enabled observability patterns using modern standards such as OpenTelemetry and a combination of open-source and commercial tooling.
- Operate as a forward deployed partner, working directly with engineering, SRE, platform teams, software development teams, and business stakeholders to solve real-world problems and deliver tailored observability implementations in production.
- Drive a transition from reactive monitoring toward proactive, intelligence-driven observability.
- Influence architectural decisions, embed observability into the software development lifecycle, guide code-level instrumentation and performance engineering practices, and ensure solutions are scalable and adaptable to diverse use cases, including event-driven and Kafka-based architectures.
- Read, troubleshoot, and guide improvements to application code; design automation, APIs, integrations, collectors, dashboards, Kafka-aware telemetry flows, and reusable tooling; and help engineering teams adopt observability as part of day-to-day software delivery.
- Translate field learnings into reusable implementation patterns, code assets, automation, Kafka/event-streaming practices, developer standards, and enterprise platform capabilities that inform platform strategy and observability standards.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field.
- Minimum of 7 years of professional experience in infrastructure engineering.
- Advanced knowledge of enterprise infrastructure technologies including cloud, network, database, storage, platform, computing, and middleware.
Qualifications
- Hands-on experience with OpenTelemetry, including application instrumentation, collector configuration, semantic conventions, and telemetry pipeline design.
- Strong software development experience in one or more languages such as Python, Go, Java, JavaScript/TypeScript, .NET/C#, Bash, or PowerShell.
- Experience with observability tools such as Prometheus, Grafana, Jaeger, Elastic, Splunk, Dynatrace, Datadog, or similar platforms.
- Strong Kafka or event streaming platform experience, including producers, consumers, topics, partitions, consumer groups, Kafka Connect, event-driven integration patterns, consumer lag analysis, throughput troubleshooting, and monitoring of Kafka-based services.
- Understanding of software development lifecycle practices including Git-based source control, code review, CI/CD pipelines, automated testing, release practices, secure coding, and developer enablement.
- Experience with infrastructure as code or configuration automation tools such as Terraform, Ansible, Helm, or similar technologies.
- Experience defining observability strategies, technical standards, developer enablement patterns, event-streaming observability patterns, and engineering roadmaps across multiple application or platform teams.
- Proven ability to translate customer-specific implementations, code assets, event-streaming patterns, and field learnings into reusable platform capabilities, enterprise standards, and strategic technology recommendations.
Skills
- Ability to read, debug, profile, and guide improvements to application code to identify performance bottlenecks, inefficient queries, memory pressure, dependency latency, threading issues, or error-handling gaps.
- Strong understanding of software development lifecycle practices including Git-based source control, code review, CI/CD pipelines, automated testing, release practices, secure coding, and developer enablement.
- Experience with infrastructure as code or configuration automation tools such as Terraform, Ansible, Helm, or similar technologies.
- Experience defining observability strategies, technical standards, developer enablement patterns, event-streaming observability patterns, and engineering roadmaps across multiple application or platform teams.
- Proven ability to translate customer-specific implementations, code assets, event-streaming patterns, and field learnings into reusable platform capabilities, enterprise standards, and strategic technology recommendations.
Benefits
All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truist’s defined benefit pension plan, restricted stock units, and/or a deferred compensation plan. As you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work. Truist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace. EEO is the Law E-Verify IER Right to Work