Jobs · Engineering · Tennessee

Principal Platform Software Engineer

Oracle · Nashville, TN · Today
Engineering$105k–$235k/yrFull-time

About the role

The Oracle Cloud Infrastructure (OCI) team offers the opportunity to build and operate massive-scale, integrated cloud services in a broadly distributed, multi-tenant cloud environment. OCI builds cloud products for customers who are solving some of the world's largest technical and business challenges. OCI Container Instances (CI) is OCI's managed serverless container offering. CI enables customers to deploy, operate and scale on OCI without the need for container orchestration, through seamless integration with OCI Compute, Networking, Storage, Identity, Security, Observability, and Developer services.

Responsibilities

  • Contribute to the Container Instances Data Plane team, designing, building, and maintaining production-grade Linux hypervisor and guest operating system images used to run containerized workloads.
  • Develop and maintain automated processes for creating, customizing, validating, and publishing Linux images in formats such as qcow2.
  • Customize Linux operating systems at the kernel, systemd, networking, storage, security and package-management layers to meet platform requirements.
  • Build and maintain RPM packages, software repositories, boot configurations and image composition workflows using DNF/YUM and, where applicable, OSTree-based technologies.
  • Develop and maintain hypervisor and lifecycle-management components in Java, Go and shell scripting.
  • Design and optimize virtualization solutions using QEMU/KVM and paravirtualized devices, including virtio-based networking and storage.
  • Profile and improve the performance, reliability, boot time, resource utilization and scalability of virtualized Linux environments.
  • Troubleshoot complex issues across the Linux kernel, virtualization stack, container runtimes, networking, storage and distributed cloud infrastructure.
  • Design and execute functional, integration, performance, regression and security testing for Linux images and virtualization components.
  • Build and enhance CI/CD pipelines that automate image creation, package integration, security validation, testing, release qualification and deployment.
  • Monitor and remediate operating system vulnerabilities by integrating security patches, kernel updates, package updates, encryption and secure-bootstrapping practices into the image lifecycle.
  • Integrate the data plane with OCI services and APIs for networking, storage, identity, telemetry and infrastructure lifecycle management.
  • Automate infrastructure provisioning and validation using Terraform and other Infrastructure as Code practices.
  • Carefully coordinate the safe rollout of new Linux image releases across hundreds of hypervisors in multiple global regions, including staged deployment, compatibility validation, monitoring and rollback planning.
  • Produce clear technical documentation, design proposals, operational procedures, and troubleshooting guides, and communicate architectural decisions and technical tradeoffs to engineering partners.

Qualifications

  • BS/MS in Computer Science, Engineering, or equivalent practical experience.
  • 6+ years of experience building and operating production software systems, preferably distributed systems or cloud services.
  • Strong system administration experience in Linux, including bash/shell scripting.
  • Strong programming experience in Java or Go.
  • Strong fundamentals in data structures, algorithms, operating systems, networking, and distributed systems.
  • Experience working with containers, Docker, Kubernetes, or related cloud-native technologies.
  • Experience with CI/CD systems, automated testing, and modern software development practices.
  • Experience using AI-assisted software development tools to improve engineering productivity across coding, testing, debugging, and documentation while maintaining high standards for code quality, security, and engineering judgment.
  • Experience using Terraform for Infrastructure as Code.
  • Strong debugging, troubleshooting, and problem-solving skills.
  • Excellent communication skills with a strong sense of ownership and the ability to collaborate effectively across teams.
  • Experience operating highly available production services, including on-call rotations, incident response, and driving operational excellence.

Similar jobs