System Engineer 3
TALENT Software Services · Redmond, WA · 1 wk ago
On-siteOTHRFull-time
About the Role
This is a site reliability engineering role in the Service Health Platforms team, responsible for the operation of enterprise network observability platforms. These platforms consist of commercially available tools and internally developed systems that ingest network telemetry, providing the foundation for tools used by network engineers to operate efficiently at scale.
The candidate will work in a large enterprise environment, gaining expertise in tying network telemetry to AI workflows for performance, capacity management, and security purposes.
Responsibilities
- Administer and operate Windows and Linux VMs hosted in Azure to ensure compliance with security and configuration standards.
- Maintain network observability platforms including syslog-ng and trapd by performing upgrades, patching, capacity planning, and authoring rules using regular expressions.
- Provide support for IBM SevOne Network Performance Manager and Broadcom AppNeta observability platforms.
- Manage assigned projects and program components to deliver services in accordance with established objectives.
- Identify and automate tasks using Bash, PowerShell, or Python.
- Troubleshoot complex network observability configurations, software applications, and operating systems through regular maintenance.
- Deploy, configure, and manage cloud services on platforms like Azure, ensuring scalability, reliability, and cost-effectiveness.
- Implement DevOps practices and tools, including CI/CD pipelines and infrastructure as code, to automate and streamline development and deployment processes.
- Perform BCDR failover testing where appropriate.
- Participate in on-call rotation (DRI).
Requirements
- Bachelor's degree in a technical field such as computer science, computer engineering, or related field; or equivalent experience.
- 5-7 years of enterprise experience in an IT systems, network, or site reliability engineering role.
- Strong understanding of network observability, including SNMP, SNMP Traps, NetFlow, and gNMI.
- Hands-on experience with cloud platforms such as Azure or similar.
- Strong understanding of enterprise networking protocols.
- Experience with system capacity and planning, as well as functional configuration and audit.
- Experience with system planning and capacity tools and analyses.
Preferred Qualifications
- Proficiency in automating tasks using Bash, PowerShell, or Python.
- Proficient with regular expressions.
- Experience with IBM SevOne and/or Broadcom AppNeta or similar observability platforms.
- Experience with various source control platforms.
- Working knowledge of Ansible and playbooks.
- Intermediate knowledge of data retrieval languages such as KQL and T-SQL.
- Experience and exposure to commercially available AI platforms.
Skills
- Syslog-NG: 3 years of experience
- Linux sysadmin: 5 years of experience
- Network Engineering: 3 years of experience
Typical Day in the Role
- Operate network observability platforms with a primary focus on syslog-ng and trapd running on Linux.
- Plan and execute security, operating system, and application patching.
- Investigate automated alerts generated from monitoring and customer-reported incidents related to network observability platforms.
- Assist network and security engineers in identifying traffic patterns and resource utilization.
- Attend standup meetings with other engineers to review and prioritize work for on-time delivery.
Performance Measurement
The contractor will be measured by compliance outcomes in internal systems such as S360, incident resolution, and the ability to deliver timely and effective outcomes for projects.
Location
Remote