Jobs · Engineering · Illinois

DevOps & Site Reliability Engineer (Digital)

Tata Consultancy Services · Deerfield, IL · 1 wk ago
Engineering$110k–$150k/yrFull-time

Skills

  • Technology and Programming (Expert Level)
    • Strong proficiency in Java full stack development
    • Object-Oriented programming principles and concepts
    • Hands-on experience with Observability platform Dynatrace
    • Hands-on experience with Spring Framework (Spring Boot, Spring MVC, Spring Security)
    • Knowledge of RESTful API development
    • Experience with databases like Oracle, DB2, MySQL
    • Proficiency in Payment Switch BASE24 EPS, C++, AS400, and Python (added advantage)
  • Domain, Cloud & Platform Engineering
    • Domain experience in Retail Point of Sale, Payment Systems, Merchandising, Inventory, or Logistics
    • Expertise in Microsoft Azure, including:
      • Compute (VMs, App Services, Azure Container Apps)
      • Containers & Orchestration (AKS, Docker)
      • Storage, Azure Key Vault, Azure Monitor, Log Analytics
    • Proven experience designing enterprise-grade, highly available cloud platforms
  • DevOps & Engineering Excellence
    • Advanced experience with Azure DevOps and CI/CD pipeline architecture
    • Strong scripting skills (PowerShell, Bash)
    • GitOps concepts, branching strategies, release orchestration
  • Site Reliability Engineering
    • Ownership of platform reliability, resiliency, and performance
    • Definition and governance of SLIs, SLOs, SLAs
    • Error budgets and reliability metrics
    • Advanced Observability Strategy, Designing, and Implementation (metrics, logs, traces, alerts, dashboards using Dynatrace)
    • Incident response leadership, RCA facilitation, and long-term remediation planning
    • Experience operating 99.9%–99.99% availability systems
  • Security, Compliance & Cost
    • Secure cloud design using Key Vault, managed identities, RBAC
    • Cost optimization (FinOps mindset) across cloud infrastructure

Responsibilities

  • Act as SRE Technical Architect for clients' Retail platforms, owning reliability and stability outcomes
  • Define and enforce SRE standards, best practices, and operating models
  • Architect and govern highly available, scalable cloud platforms
  • Lead the design and implementation of CI/CD and IaC strategies
  • Establish proactive monitoring, alerting, and incident prevention mechanisms
  • Own major incident leadership, RCA execution, and corrective action tracking
  • Partner with application, security, and architecture teams to build reliability by design
  • Drive automation to reduce toil and improve operational efficiency
  • Mentor and coach SRE and DevOps engineers across teams
  • Influence roadmap decisions with a reliability, scalability, and cost lens

Benefits

  • Discretionary Annual Incentive
  • Comprehensive Medical Coverage:
    • Medical & Health, Dental & Vision
    • Disability Planning & Insurance
    • Pet Insurance Plans
  • Family Support:
    • Maternal & Parental Leaves
  • Insurance Options:
    • Auto & Home Insurance
    • Identity Theft Protection
  • Convenience & Professional Growth:
    • Commuter Benefits
    • Certification & Training Reimbursement
  • Time Off:
    • Vacation, Time Off, Sick Leave & Holidays
  • Legal & Financial Assistance:
    • Legal Assistance
    • 401K Plan
    • Performance Bonus
    • College Fund
    • Student Loan Refinancing

Pay

Salary Range: $110,000–$150,000 a year

Similar jobs