Site Reliability Engineer - VP
Barclays · New York, NY · Today
Engineering$170k/yrFull-time
About the role
To manage the IT Services department and set the strategic direction, provide support to the bank's senior management team, and to manage IT Service risk across the organisation.
Responsibilities
- Development of strategic direction for IT Services, including the implementation of up-to-date methodologies and processes.
- Management of the IT Services department, including oversight of IT Services colleagues and their performance, implementation of departmental goals and objectives, oversight of department efficiency and effectiveness.
- Relationship management of IT Services stakeholders, including identifying relevant stakeholders, and maintenance of the quality of external third party services.
- Development and implementation of policies and procedures for IT Services, implementation and adherence of control targets and standards, policies and procedures for IT Services, managing adherence to group SLAs and controls associated with core technology production activities in incident, problem, and change.
- Monitoring the financial performance of the IT Services department, including revenue, profitability, and cost control, driving value from any commercial agreements, strong management of any directly controlled costs etc.
- Management of IT Services projects, including driving successful research and related product launches, and deliverance of integrated solutions to clients.
- Effectively monitor and maintain the bank’s critical technology infrastructure and resolve more complex technical issues, whilst minimising disruption to operations.
Requirements
- Good SRE experience with experience implementing reliability practices, defining SLOs/error budgets, uptime and supporting incident response for multi-layered distributed systems.
- Proficient in cross region cloud platforms, AWS, Azure, GCP, observability tools, Prometheus and CloudWatch, and automation frameworks with experience of reducing toil.
- Solid infrastructure as code skills, Terraform and CloudFormation with experience building self-service platforms, tooling, and deployment pipelines for engineering teams.
- Demonstrated ability in capacity planning, performance tuning, and cost optimization with experience managing production systems at significant scale.
- Experience in observability and monitoring, metrics collection, logging, distributed tracing, and dashboard creation.
Qualifications
- Understanding of Large Language Models and Agentic frameworks.
- Self-starter with leadership experience mentoring SREs, running post-mortems, and driving reliability improvements across engineering teams with good collaboration skills.
- Experience with Kubernetes, service mesh technologies, and microservices reliability patterns in production environments.
- Experience working with REST API development and Integration, Database management / Query language.
- Self-motivated, resourceful, and highly adaptable, with the assurance to work independently while also contributing effectively within a collaborative team environment.
Skills
- Experience with cross region cloud platforms, AWS, Azure, GCP.
- Proficiency in observability tools, Prometheus and CloudWatch.
- Strong infrastructure as code skills, Terraform and CloudFormation.
- Experience in capacity planning, performance tuning, and cost optimization.
- Knowledge of large language models and agentic frameworks.
- Leadership experience mentoring SREs, running post-mortems, and driving reliability improvements.
- Experience with Kubernetes, service mesh technologies, and microservices reliability patterns.
- Experience with REST API development and Integration, Database management / Query language.
- Self-motivation, resourcefulness, and adaptability.
Pay
Minimum Salary: $170,000
Maximum Salary: $230,000
Schedule
The role is located in our New York, NY office.