Cloud Engineer - Linux and Automation
Thrive · United States · 1 mo ago
RemoteRemoteInformation TechnologyFull-time
Responsibilities
- Administer, monitor, and troubleshoot Linux systems (primarily Ubuntu) across physical and virtual environments.
- Serve as a Tier 2 escalation point for infrastructure incidents; perform root cause analysis and implement remediation actions.
- Manage and operate virtual machine workloads within HPE Morpheus Enterprise HVM and VMware ESXi, including VM provisioning, lifecycle management, and hypervisor-level troubleshooting.
- Define and enforce standardized build documentation
- Develop and maintain Ansible playbooks for configuration management, OS patching, application deployment, and compliance enforcement.
- Write and manage Terraform configurations for infrastructure provisioning and lifecycle management.
- Enforce infrastructure-as-code best practices including version control and peer review.
- Maintain and update NetBox as the authoritative source of truth for IP address management (IPAM), DCIM asset records, and network topology documentation.
- Integrate Netbox with IaC tools for automation of routine physical equipment commissioning.
- Participate in change management processes, ensuring all changes are documented, reviewed, and approved in alignment with Thrive's ITSM standards.
- Perform routine system health checks, capacity reviews, and performance tuning for Linux servers.
- Maintain and improve internal runbooks, technical documentation, and knowledge base articles.
- Collaborate with networking, security, and cloud teams to resolve cross-functional infrastructure issues.
- Participate in an on-call rotation to support critical infrastructure incidents outside of standard business hours.
- Identify opportunities to improve operational efficiency through automation and standardization.
Requirements
- 3-5+ years of hands-on Linux system administration experience with strong proficiency in Ubuntu Linux including systemd, networking, storage management, and security hardening
- 2+ years of experience with Infrastructure as Code tools, specifically Ansible (roles, playbooks, inventories) and Terraform (state management, CI/CD plan/apply workflows)
- Hands-on familiarity with HPE Morpheus Enterprise or similar HVM/KVM platforms for VM provisioning, template management, and self-service automation; experience with VMware vCenter is a strong plus
- Working knowledge of NetBox for IPAM, DCIM, and network source-of-truth management; experience with NetBox API integrations preferred
- Proficiency in scripting languages for automation and tooling development
- Solid understanding of networking fundamentals including TCP/IP, VLANs, DNS, DHCP, routing, and firewall rules as they apply to virtualized and cloud environments
- Familiarity with Git-based version control and collaborative workflows (GitLab or GitHub)
- Able to diagnose and resolve complex system and infrastructure issues independently and under pressure
- Strong documentation habits - able to produce accurate and maintainable runbooks, topology diagrams, and change records
- Strong team player - collaborates effectively with peers and cross-functional teams to solve problems
- Proactive, change-oriented mindset - actively seeks process improvements and drives automation-first approaches
- Ability to communicate technical concepts clearly to both technical and non-technical stakeholders
Qualifications
- Bachelor's degree in Computer Science, or a related discipline — or equivalent combination of education and relevant work experience
- Knowledge of ITIL and ITSM best practices
Benefits
- Work Schedule: Standard business hours with participation in an on-call rotation.
- Remote work eligible
- Occasional travel may be required for datacenter activities.
- Less than 10%
Pay
Salaries are determined according to the job's scope, market data, location, and the candidate’s qualifications, including experience and relevant education.
Schedule
Standard business hours with participation in an on-call rotation.
Remote work eligible
Occasional travel may be required for datacenter activities.
Less than 10%