SRE Cyber Tools - Splunk Admin.
About the Role
Leidos is seeking a Site Reliability Engineer (SRE) Splunk Enterprise Administrator focused on Cyber support for the largest IT services program for the Navy. Under the Service Management, Integration, and Transport (SMIT) program, the Leidos team delivers the core backbone of the Navy-Marine Corps Intranet, including cybersecurity services, network operations, service desk, and data transport. Leidos supports the Navy in unifying its shore-based networks and data management to improve capability and service while reducing costs through a consolidated enterprise network.
As part of the SRE organization, you will develop and execute tests focused on system resilience, performance under load, and failure scenarios. You will collaborate with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions, validating system behavior under normal and stress conditions to ensure services are resilient and meet established Service Level Objectives (SLOs). The SRE will support the operations and maintenance of the enterprise network, contributing to the development of robust and scalable services that operate reliably in production.
The Splunk Enterprise Administrator supports mission-critical cybersecurity operations by administering and maintaining distributed Splunk Enterprise platforms across hybrid cloud and on-premises environments. The role involves daily platform operations, performance and ingestion monitoring, data onboarding support, incident troubleshooting, security compliance, documentation, and modernization activities in an Agile/DevOps operating culture.
Key metrics of success for the team include:
- Improved system reliability, as measured by adherence to Service Level Objectives (SLOs) and reduced Mean Time to Recovery (MTTR).
- Comprehensive and regularly updated automated test coverage for all critical systems and infrastructure components.
- Timely identification and resolution of performance bottlenecks and failure points.
- Integration of automated testing into the CI/CD pipeline, ensuring continuous reliability validation.
- Increased scalability and performance of systems under high load due to effective performance testing.
Responsibilities
- Administer, maintain, and perform daily Operations and Maintenance (O&M) for distributed Splunk Enterprise environments, including Search Heads, Indexers, Heavy Forwarders, Intermediate Forwarders, Universal Forwarders, and Deployment Servers across hybrid cloud and on-premises infrastructure.
- Monitor platform health, data ingestion, indexing throughput, search performance, and retention utilization; troubleshoot data onboarding, parsing, field extraction, forwarding, indexing, and ingestion issues.
- Support Splunk Cloud integrations and associated hybrid operational activities.
- Maintain Splunk applications, dashboards, alerts, saved searches, and knowledge objects.
- Support patching, vulnerability remediation, Security Technical Implementation Guide (STIG) compliance, and system hardening activities.
- Support platform upgrades, migrations, infrastructure modernization, technology refresh, and automation initiatives using scripting and configuration-management tools.
- Participate in incident response, outage troubleshooting, problem resolution, and root cause analysis.
- Maintain operational documentation, architecture diagrams, runbooks, and standard operating procedures (SOPs), and coordinate with cybersecurity, network, server, engineering, and customer teams during operational and modernization activities.
- Participate in after-hours support and an on-call rotation as required.
Requirements
- Requires a BS degree and 5-10 years of prior relevant experience or a Master’s with 4-8 years of prior relevant experience.
- U.S. Citizen and possess an active Secret Security Clearance.
- Minimum of DoD 8570.01 IAT Level II Certification required.
- Minimum three years of experience administering Splunk Enterprise v9 and supporting distributed Splunk architectures in hybrid cloud and on-premises production environments.
- Strong working knowledge of Splunk data ingestion pipelines, indexing, parsing, forwarding, search optimization, and retention management.
- Experience designing, developing, and maintaining operational dashboards, visualizations, and executive reporting in Splunk Enterprise and Splunk Cloud, including IT Service Intelligence (ITSI), Service Analyzer, Glass Tables, and KPI-driven service health monitoring.
- Experience administering and troubleshooting Red Hat Enterprise Linux (RHEL) 8 and/or RHEL 9 servers.
- Working knowledge of cybersecurity monitoring and Security Information and Event Management (SIEM) operations, with demonstrated experience troubleshooting operational incidents in complex enterprise environments.
- Working knowledge of TCP/UDP networking, SSL certificates, Syslog, REST APIs, and authentication integrations.
- Experience using at least one scripting or automation language: Python, Bash, or PowerShell.
- Strong troubleshooting, analytical, communication, and documentation skills, with the ability to work independently, manage competing priorities, collaborate across technical teams, and maintain a customer-focused approach in a high-tempo production environment.
- Must have a vendor certification, e.g., Splunk Enterprise Certified Admin, Splunk Cloud Certified Admin, or Scaled Agile Framework (SaFe).
- Ability to work onsite at Norfolk Naval Station Monday through Friday day shift.
Preferred Qualifications
- Splunk Core Certified Power User and/or Splunk Enterprise Certified Admin certification.
- Experience supporting Splunk Cloud environments or integrations and both Splunk Enterprise v9 and v10.
- Experience working in classified government or Department of Defense environments.
- Familiarity with STIGs, the Risk Management Framework (RMF), vulnerability management, and compliance frameworks.
- Familiarity with RHEL 10, DevOps practices, and Infrastructure-as-Code concepts.
- Experience with automation or configuration-management tools such as Ansible, Jenkins, Chef, or Terraform.
Pay
Pay Range: $107,900.00 - $195,050.00. The Leidos pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include (but are not limited to) responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.
Benefits
Pay and benefits are fundamental to any career decision. Employment benefits include competitive compensation, Health and Wellness programs, Income Protection, Paid Leave, and Retirement. More details are available at www.leidos.com/careers/pay-benefits.