Database Site Reliability Engineer
We are seeking an experienced Database Site Reliability Engineer (SRE) to support and operate mission-critical database platforms within a fast-paced enterprise environment. This role is focused on operational excellence, reliability, resiliency, automation, and continuous improvement across multiple database technologies. This is not a development role. We are looking for hands-on database professionals who thrive in production operations, take ownership of issues, understand urgency, and are passionate about improving systems, processes, and themselves. The ideal candidate views reliability as a product, proactively identifies risks before they become incidents, and continuously seeks opportunities to automate repetitive tasks and improve platform stability.
About the role
This is a contract assignment with one of our premier Financial Services clients in Alpharetta, GA. The shifts will be 4-10 hour days (Sunday-Wednesday 7am-5pm).
Responsibilities
- Database Operations & Reliability
- Install, configure, upgrade, patch, and maintain enterprise database platforms.
- Ensure availability, performance, recoverability, and security of production database environments.
- Monitor database platforms and associated infrastructure, responding rapidly to incidents and service degradations.
- Lead troubleshooting efforts for database, operating system, storage, replication, and application connectivity issues.
- Execute failovers, disaster recovery testing, and recovery procedures.
- Partner with application teams to provide database guidance and operational support.
- Platform Engineering
- Deploy, maintain, and optimize database infrastructure across physical, virtual, and cloud environments.
- Implement scalable, resilient database solutions.
- Evaluate and recommend improvements to architecture, monitoring, automation, and operational processes.
- Support capacity planning, performance tuning, and platform lifecycle management.
- Automation & Continuous Improvement
- Develop and maintain automation solutions using Python, Shell, Ansible, or similar technologies.
- Help eliminate manual operational activities through engineering and automation.
- Improve monitoring, alerting, reporting, and operational workflows.
- Drive incremental improvements that reduce risk, improve reliability, and increase operational efficiency.
- Performance & Incident Management
- Analyze and resolve database performance issues.
- Troubleshoot replication, backup/recovery, storage, network, and infrastructure-related incidents.
- Participate in root cause analysis and drive permanent corrective actions.
- Review operational metrics and trends to identify opportunities for improvement.
- Operational Excellence
- Maintain accurate operational documentation, standards, and procedures.
- Generate and present operational metrics, service health indicators, and reliability reporting.
- Participate in incident response activities.
- Demonstrate strong ownership from issue identification through resolution.
Requirements
- Strong experience administering enterprise database platforms, including:
- Sybase ASE
- Oracle RAC
- Additional database technologies such as MongoDB, Cassandra, Redis, PostgreSQL, MySQL, or similar platforms are a plus.
- Experience performing:
- Installation
- Configuration
- Upgrades
- Patching
- Performance tuning
- Backup and recovery
- High availability and disaster recovery
- Experience with database replication technologies including:
- SAP Replication Server
- Data Guard
- HVR (preferred)
- Strong Linux administration skills.
- Experience with automation and scripting:
- Python
- Ansible
- Shell scripting
- Understanding of storage, networking, operating systems, and infrastructure services.
- Experience with Veritas Cluster Server, ASM, LVM, and SAN technologies.
- Familiarity with enterprise operational tooling such as Jira, Service Now, and Confluence.
- Strong analytical, troubleshooting, and problem-solving skills.
Schedule
4-10 hour days (Sunday-Wednesday 7am-5pm).