DevOps & Site Reliability Engineer (Digital)
Tata Consultancy Services · Deerfield, IL · 1 wk ago
Engineering$110k–$150k/yrFull-time
Skills
- Technology and Programming (Expert Level)
- Strong proficiency in Java full stack development
- Object-Oriented programming principles and concepts
- Hands-on experience with Observability platform Dynatrace
- Hands-on experience with Spring Framework (Spring Boot, Spring MVC, Spring Security)
- Knowledge of RESTful API development
- Experience with databases like Oracle, DB2, MySQL
- Proficiency in Payment Switch BASE24 EPS, C++, AS400, and Python (added advantage)
- Domain, Cloud & Platform Engineering
- Domain experience in Retail Point of Sale, Payment Systems, Merchandising, Inventory, or Logistics
- Expertise in Microsoft Azure, including:
- Compute (VMs, App Services, Azure Container Apps)
- Containers & Orchestration (AKS, Docker)
- Storage, Azure Key Vault, Azure Monitor, Log Analytics
- Proven experience designing enterprise-grade, highly available cloud platforms
- DevOps & Engineering Excellence
- Advanced experience with Azure DevOps and CI/CD pipeline architecture
- Strong scripting skills (PowerShell, Bash)
- GitOps concepts, branching strategies, release orchestration
- Site Reliability Engineering
- Ownership of platform reliability, resiliency, and performance
- Definition and governance of SLIs, SLOs, SLAs
- Error budgets and reliability metrics
- Advanced Observability Strategy, Designing, and Implementation (metrics, logs, traces, alerts, dashboards using Dynatrace)
- Incident response leadership, RCA facilitation, and long-term remediation planning
- Experience operating 99.9%–99.99% availability systems
- Security, Compliance & Cost
- Secure cloud design using Key Vault, managed identities, RBAC
- Cost optimization (FinOps mindset) across cloud infrastructure
Responsibilities
- Act as SRE Technical Architect for clients' Retail platforms, owning reliability and stability outcomes
- Define and enforce SRE standards, best practices, and operating models
- Architect and govern highly available, scalable cloud platforms
- Lead the design and implementation of CI/CD and IaC strategies
- Establish proactive monitoring, alerting, and incident prevention mechanisms
- Own major incident leadership, RCA execution, and corrective action tracking
- Partner with application, security, and architecture teams to build reliability by design
- Drive automation to reduce toil and improve operational efficiency
- Mentor and coach SRE and DevOps engineers across teams
- Influence roadmap decisions with a reliability, scalability, and cost lens
Benefits
- Discretionary Annual Incentive
- Comprehensive Medical Coverage:
- Medical & Health, Dental & Vision
- Disability Planning & Insurance
- Pet Insurance Plans
- Family Support:
- Maternal & Parental Leaves
- Insurance Options:
- Auto & Home Insurance
- Identity Theft Protection
- Convenience & Professional Growth:
- Commuter Benefits
- Certification & Training Reimbursement
- Time Off:
- Vacation, Time Off, Sick Leave & Holidays
- Legal & Financial Assistance:
- Legal Assistance
- 401K Plan
- Performance Bonus
- College Fund
- Student Loan Refinancing
Pay
Salary Range: $110,000–$150,000 a year