Jobs · Information Technology · Washington

Senior Core Infrastructure Engineer (OCI Object Storage)

hackajob · Seattle, WA · 2 wk ago
On-siteInformation Technology$79k–$210k/yrFull-time

About the role

Oracle's Cloud Infrastructure team is building Infrastructure-as-a-Service technologies that operate at high scale in a broadly distributed multi-tenant cloud environment. The Object Storage Service team is looking for hands-on engineers with expertise in solving difficult problems in distributed systems, large-scale storage, and highly available services. As a senior engineer, you will own the software design and development for major components and features of the Object Storage Service.

Responsibilities

  • System Scalability – Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools; optimizes code and/or systems for large-scale data processing; implements scalability requirements for assigned components; leverages components of data plane platforms to handle large-scale data retrieval, storage, and processing; implements performance and load testing
  • System Reliability Design – Collaborates with team to build fault-tolerant components capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms; applies recovery oriented computing principles to design components that effectively handle service disruptions; implements retry mechanisms, circuit breakers, and timeouts to help handle network unreliability
  • System Reliability Performance – Implements tests and alarm configurations to proactively detect and address issues/failures; supports efforts to recover from failures by drafting and executing runbooks and operational procedures; builds and customizes dashboards, telemetry systems, and alerting mechanisms to monitor component health
  • Correctness / Availability – Designs and implements functional requirements and testing for assigned features within an existing system; implements tests scenarios (e.g., fault-injection, brown-out) to evaluate system correctness; implements standard data replication and synchronization techniques to maintain data integrity and availability
  • Operational Troubleshooting & Incident Management – Diagnoses, debugs, and resolves issues in system components to support ongoing operation; implements strategies to prevent interruptions, ensuring no maintenance windows are required; designs and implements automation scripts and tooling used to troubleshoot operational issues; participates in operational support rotations, assisting in incident responses and root cause investigations
  • Compliance & Security – Applies advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls; implements remediation plans to continuously improve security; collaborates to ensure cloud infrastructure complies with relevant industry standards and regulations
  • Automation & Change Management – Maintains automation scripts and tools (e.g., Infrastructure as Code) for managing cloud infrastructure; adheres to change management plans for patching, updating, and rolling back applications
  • Planning & Execution – Tracks timelines with minimal supervision, ensuring work is completed in a timely manner and in alignment with project requirements; prioritizes and adjusts work as resources or timelines change with some guidance
  • Collaboration & Partnership – Collaborates across teams to align on expectations and achieve shared objectives; builds and maintains comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships; actively listens to diverse perspectives and asks questions to ensure understanding of others
  • Problem Solving – Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate; analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors; contributes to knowledge sharing and best practices
  • Continuous Learning – Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools, and staying current with industry trends and best practices; seeks out and leverages feedback and training to improve skills; contributes to a culture of continuous learning and knowledge sharing with team members
  • Continuous Improvement – Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team; seeks input from team members on alternative approaches and methods for improving work

Qualifications

  • Rock-solid coder with familiarity with distributed systems
  • Values simplicity and scale
  • Works comfortably in a collaborative, agile environment
  • Excited to learn

Similar jobs