Site Reliability Engineer
Canonical · NAMER · 3 days ago
RemoteRemoteEngineeringFull-time
Job Summary
The role involves deploying and running OpenStack, Kubernetes, storage solutions, and open source applications, applying DevOps practices. The ideal candidate should have a strong background in Linux, Python, networking, and knowledge of how clouds work.
About the Role
We are looking for a Site Reliability Engineer to join our team. This role is part of a pioneering tech firm at the forefront of the global move to open source. The company is a leader in providing open source software and operating systems to the global enterprise and technology markets.
Responsibilities
- Deploy and run OpenStack, Kubernetes, storage solutions, and open source applications
- Identify and address incidents, monitor and observe applications
- Anticipate potential issues and enable product refinement
- Work on the entire stack, from bare-metal networking and kernel up to Kubernetes and open source applications
- Train in our core technologies like OpenStack, Kubernetes, security standards, open source products like Kubeflow, Kafka, OpenSearch, databases, and many others
- Approach automation with a scientific mindset to bring operations at scale, driven by metrics and code
Requirements
- Degree in software engineering or computer science
- Python software development experience
- Operational experience in Linux environments
- Experience with Kubernetes deployment or operations
Qualifications
- Genuine interest in the full open source infrastructure stack from bare metal to containers
- Ability to work in operations with mission-critical services for global brand-name customers
- Experience working in a distributed work environment
Skills
- Familiarity with OpenStack deployment or operations
- Familiarity with public cloud deployment or operations
- Familiarity with private cloud management
Benefits
- Colleague benefits including personal learning and development budget, annual holiday leave, maternity and paternity leave, Employee Assistance Programs, and opportunity to travel to new locations to meet colleagues
- Compensation review every 6 months to recognize outstanding performance
- Recognition rewards
- Priority Pass and travel upgrades for long-haul company events
What We Are Looking For In You
- Fluent in Python
- Interest in the full open source infrastructure stack from bare metal to containers
- Able to work in operations with mission-critical services for global brand-name customers
- Ability to travel internationally twice a year, for company events up to two weeks long
Location
This is a globally remote role.