Jobs · Engineering

AI Native, Tech Ops Engineer

ConsumerAffairs · United States · 1 mo ago
RemoteRemoteEngineeringFull-time

Responsibilities

  • Monitor and maintain the organization's infrastructure, including servers, networks, storage systems, and applications
  • Perform routine system checks and preventive maintenance to ensure optimal performance and uptime
  • Respond to system alerts and incidents, diagnosing and resolving issues promptly to minimize downtime
  • Provide technical support to resolve infrastructure-related issues, working closely with other technical teams
  • Troubleshoot and resolve hardware, software, and network issues, escalating to higher-level support when necessary
  • Maintain detailed documentation of issues, solutions, and processes to improve the team's knowledge base
  • Identify opportunities to automate routine tasks and processes, improving operational efficiency and reducing manual workload
  • Implement scripts, automation tools, and AI skills to streamline system management and monitoring
  • Continuously evaluate and optimize infrastructure performance, capacity, and resource utilization
  • Support the development and execution of disaster recovery plans to ensure business continuity in case of system failures
  • Manage backup and restore processes for critical systems and data, ensuring data integrity and availability
  • Participate in regular disaster recovery testing and drills
  • Plan and execute decommissioning of legacy infrastructure, including EC2 instances, VPCs, and load balancers, coordinating Terraform state cleanup and DNS cutover
  • Collaborate with development, network, and security teams to ensure alignment and effective communication on infrastructure projects
  • Communicate effectively with non-technical stakeholders, providing updates on system status and issues

Requirements

  • Minimum Qualifications & Credentials: Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent work experience. 5+ years of experience in system administration, or a similar role. 5+ years of professional experience in Linux administration, managing AWS resources, developing CI/CD and server orchestration pipelines, scripting and monitoring.
  • Hard/Technical Skills: Expert in Cloud-based production systems at scale Amazon Web Services (EC2, VPC, EFS, S3, EKS etc.). Production experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments. Infrastructure as Code tools, primarily Terraform. Working in a Python and JavaScript-centric codebase and are familiar with their related best-practices. Creating CI/CD pipelines with Jenkins, Concourse or other CI/CD implementation. Monitoring tools, like Datadog or Prometheus. Scripting for server side automation, auditing, and monitoring. Experience maintaining logging, monitoring, and alerting capabilities using OpenSearch, Vector log pipelines, Prometheus, and Kafka. Configuring and managing data sources like PostgreSQL/Aurora RDS, OpenSearch, Redis, and message streaming platforms like Kafka. Design and maintain log ingestion pipelines (e.g., Vector → OpenSearch, Vector → Kafka), including index retention, document shape optimization, and failure recovery. Triage and remediate security vulnerabilities (CVEs) across infrastructure components, including container base images, OS packages, and third-party services.
  • The ideal candidate would also have: A working knowledge of modern software practices and technologies such as Agile methodologies. Promoting and establishing development standard methodologies for AWS infrastructure-as-code. Experience with AWS Well-Architected principles. Experience with High Availability implementations. Experience around Security and Compliance.

Benefits

  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off (Vacation, Sick & Public Holidays)
  • Family Leave (Maternity, Paternity)
  • Short Term & Long Term Disability
  • Training & Development

Similar jobs

AI Ops Engineer

HaystackWashington, VA· 3 wk ago
RemoteEngineeringapply on haystack.cv

AI Ops Engineer

EndavaUnited States· 1 mo ago
RemoteEngineeringapply on jobs.smartrecruiters.com

AI Ops Engineer

EndavaColorado, United States· 1 mo ago
apply on jobs.smartrecruiters.com

AI Ops Engineer

Tata Consultancy ServicesEden Prairie, MN· 1 mo ago
Engineeringapply on ibegin.tcsapps.com