Jobs · Information Technology · North Carolina

Observability Lead-Sr. Infrastructure Engineer

Truist · Charlotte, NC · 4 days ago
Information TechnologyFull-time

About the role

We are seeking a highly experienced and strategic Observability Lead-Sr. Infrastructure Engineer with a strong Forward Deployed Engineering (FDE) focus to define, implement, and scale enterprise observability capabilities.

Responsibilities

  • Establish the vision and direction for metrics, logs, traces, and telemetry pipelines using modern standards such as OpenTelemetry and a combination of open-source and commercial tooling.
  • Operate as a forward deployed partner, working directly with engineering, SRE, platform teams, and business stakeholders to solve real-world problems and deliver tailored observability implementations in production.
  • Drive a transition from reactive monitoring toward proactive, intelligence-driven observability.
  • Influence architectural decisions, embed observability into the software development lifecycle, and ensure solutions are both scalable and adaptable to diverse use cases.
  • Improve system reliability, reduce mean-time-to-detect and resolve (MTTD/MTTR), enable faster root-cause analysis, and create a consistent, high-fidelity observability experience.
  • Translate field learnings into reusable patterns that inform platform strategy and enterprise standards.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field.
  • Minimum of 7 years of professional experience in infrastructure engineering.
  • Advanced knowledge of enterprise infrastructure technologies including cloud, network, database, storage, platform, computing, and middleware.

Qualifications

  • Hands-on experience with OpenTelemetry, including instrumentation, collector configuration, and pipeline design.
  • Experience with observability tools such as Prometheus, Grafana, Jaeger, Elastic, Splunk, Dynatrace, or similar platforms.
  • Strong background in Kubernetes, cloud platforms, and service-oriented or event-driven architectures.
  • Proficiency in one or more programming or scripting languages (e.g., Python, Go, Java, Bash) for automation and integration.
  • Prior experience in forward deployed engineering, solutions engineering, or customer-facing technical roles.
  • Proven ability to translate customer-specific implementations into reusable platform capabilities and enterprise standards.
  • Demonstrated success in defining observability strategies and influencing platform and engineering roadmaps.

Skills

  • Strong technical skills in infrastructure engineering.
  • Knowledge of modern observability tools and standards.
  • Ability to collaborate effectively with cross-functional teams and external partners.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication and stakeholder management skills.

Benefits

All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truist’s defined benefit pension plan, restricted stock units, and/or a deferred compensation plan.

Similar jobs