Jobs · Information Technology · California

Senior Software Engineer, Datacenter Operations Platform Engineering

Crusoe · San Francisco, CA · 2 days ago
On-siteInformation TechnologyFull-time
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster. We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI. We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. About the Role: We are seeking Senior Software Engineers to design and develop internal datacenter tooling and infrastructure management systems for Crusoe Cloud, a leading cloud provider. You will play a crucial role in building tools for datacenter facilities managers and capacity planners, as well as creating automation software to efficiently bring server hardware, switches, and other infrastructure components online. You’ll help evaluate and implement tools and frameworks for our internal customers, focusing on reliability, scalability, operational efficiency, and ease of use. This role is central to streamlining infrastructure management processes and enhancing our cloud platform’s overall performance as we dramatically scale our hardware footprint. Blog Posts About This Team: Before first boot: How Crusoe Pre-Deployment Automation readies every GPU serverHealthy by design: How Crusoe burn-in tests every node before it reaches you What You’ll Be Working On: Designing and developing advanced internal tooling for datacenter facilities managers and capacity plannersCreating automation software for rapid deployment and configuration of servers, network switches, power delivery units (PDUs), and coolant delivery units (CDUs)Mentoring junior engineers through design guidance and code reviews to ensure high-quality solutionsInnovating and implementing features that streamline infrastructure management and operational capabilitiesCollaborating with cloud support and operations teams to develop tools that enable growth and empower internal processesPartnering cross-functionally to align goals and optimize resource utilization for improved infrastructure managementLeading by example in technical excellence and fostering an environment of innovation in infrastructure and tooling development What You’ll Bring to the Team: 5+ years of professional software development experience5+ years of programming experience in at least one modern compiled language (Go, Rust, Java, or C++)5+ years of experience contributing to architecture and design (patterns, reliability, scaling) of new and existing systemsBachelor’s degree in Computer Science or related field, or 5–8+ years of equivalent experienceStrong computer science fundamentals in data structures and algorithmsProven experience building and maintaining scalable, highly available, fault-tolerant distributed systemsSolid understanding of infrastructure design and operational trade-offsFamiliarity with CI/CD practices and build systems (GitLab CI/CD, CircleCI, GitHub Actions)Familiarity with modern infrastructure tools (Docker, Kubernetes, Ansible, CloudFormation, Terraform)Experience with concurrency, multithreading, and synchronizationExperience with Unix/Linux environmentsExperience with TCP/IP and network programmingExcellent communication skillsAlignment with company values Bonus Points: Experience working in large-scale datacenter or cloud environmentsHands-on exposure to GPU clusters or high-performance computing environmentsBackground in infrastructure observability or monitoring systems Benefits: Industry competitive payRestricted Stock Units in a fast growing, well-funded technology companyHealth insurance package options that include HDHP and PPO, vision, and dental for you and your dependentsEmployer contributions to HSA accountsPaid Parental LeavePaid life insurance, short-term and long-term disabilityTeladoc401(k) with a 100% match up to 4% of salaryGenerous paid time off and holiday scheduleCell phone reimbursementTuition reimbursementSubscription to the Calm appMetLife LegalCompany paid commuter benefit; $300 per month Compensation: Compensation will be paid in the range of $170,000 - $205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data. Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Similar jobs