Site Reliability Engineer Intern
Copart · Dallas, TX · 1 wk ago
EngineeringFull-time
Copart, Inc., a technology leader and the premier online vehicle auction platform globally, links vehicle sellers to more than 750,000 buyers in over 190 countries. We believe in providing an unmatched experience driven by our people, processes, and technology.
About the role
We are seeking a highly skilled and proactive Site Reliability Engineer Intern to join our SRE team in Dallas. This role is pivotal in ensuring the stability, performance, and reliability of Copart’s critical applications and global data center infrastructure through advanced monitoring, troubleshooting, and automation. As a key member of our 24/7 operations, you will contribute to meeting our stringent SLA commitments and play a crucial role in maintaining operational excellence.
Responsibilities
- Proactively monitor Copart’s global data centers and application infrastructure using internal tools, identifying and resolving issues before they impact operations.
- Design, build, and optimize monitoring and automation tools using Python and Ansible to collect key metrics and enable automated remediation of infrastructure and application issues.
- Maintain and optimize a suite of monitoring and collaboration tools, including Datadog and Kubernetes-related observability tooling.
- Conduct monthly security patching for operating systems and critical applications.
- Support Infrastructure teams by troubleshooting and communicating issues across various Copart domains.
- Partner with cross-functional internal teams, including Product Development, DevOps, Network, Systems, and Database teams.
- Develop analytical and reporting capabilities to monitor performance, identify areas for improvement, and implement quality control plans using Datadog.
- Create and maintain comprehensive standard operating procedures (SOPs), system diagrams, and training materials for team use.
- Support incident management activities including triage, tracking, root cause analysis, and monthly review of lessons learned.
- Process change paperwork throughout the day as part of daily operational requirements.
- Coordinate and perform periodic failover testing for Copart’s network, systems infrastructure, and application environments to ensure business continuity.
- Participate in root cause analysis and post-incident review processes to identify trends, recurring issues, and opportunities for preventive action.
Requirements
- Experience triaging, tracking, and resolving incidents in a production environment.
- Ability to perform initial impact assessment and severity classification.
- Strong communication skills for updates to stakeholders during active incidents.
- Experience documenting incident timelines, actions taken, and resolution details.
Qualifications
- Collaborative & Communicative: A strong team player with excellent interpersonal, oral, and written communication skills who thrives in a collaborative environment.
- Technically Versatile: Core knowledge of Linux and Windows, with experience in virtual environments, basic networking, scripting, automation, and observability tools. Familiarity with Datadog is a plus.
- Innovative & Proactive: Continuously looks for ways to improve processes and procedures and is comfortable suggesting and implementing practical solutions.
- Self-Starter & Adaptable: Highly motivated individual who can work independently, manage competing priorities, and perform effectively in a flexible work schedule.
Skills
- Intermediate proficiency in programming and scripting.
- Strong troubleshooting skills across Linux/Unix/Windows-based systems.
- Hands-on experience with virtual machine management software, particularly VMware vSphere (a plus).
- Familiarity with monitoring and observability tools such as Datadog.
- Basic understanding of AI tools and concepts, including AI-assisted workflows, prompt-based tools, or automation-enhanced support processes.
- Ability to leverage AI tools to improve efficiency in troubleshooting, documentation, analysis, or operational support (a plus).
- Experience with AWS or GCP (a plus).