Mid Systems Engineer (w/ active Secret)
CriticalSolutions, LLC · Norfolk, VA · 3 wk ago
$85k–$110k/yrFull-time
Location: Norfolk, VA (also available in San Diego, CA, Pearl Harbor, HI, or Bremerton, WA). Must possess an active DoD Secret Clearance and be able to pass a background investigation and fingerprinting. US citizenship is required.
Responsibilities
- Work alongside development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall platform reliability.
- Develop features using AI coding tools and a repository of scripts to automate, scale, test, and secure cloud infrastructure and pipelines.
- Enhance performance monitoring of systems via Splunk or other dashboard reporting tools.
- Identify performance bottlenecks and optimize cloud infrastructure performance.
- Contribute to the SRE journey by suggesting improvements for engineering build, maintenance, automation, and reliability using SRE/DevOps tools and Infrastructure-as-Code.
- Develop and code high-quality pipeline automation workflows supporting business and technology strategies both inside and outside the cloud platform.
- Develop and execute test strategies simulating real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
- Create, script, and run performance tests to measure system behavior under varying loads, identify bottlenecks, and optimize performance.
- Design, implement, and maintain automated test suites for infrastructure and application components, integrating testing into the CI/CD pipeline.
- Build automated systems for continuous performance, stress, and load testing.
- Collaborate with SREs, developers, and operations teams to define reliability goals and develop testing strategies to validate them.
- Ensure new services and features undergo thorough testing for performance, reliability, and failure recovery before production deployment.
- Validate monitoring, logging, and alerting mechanisms by testing systems under failure conditions.
- Ensure Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks.
- Resolve conflicts between timeline, budget, and scope independently; escalate sophisticated or consequential issues to senior management.
- Design, deploy, and maintain Microsoft Endpoint Configuration Manager (MECM) environments at enterprise scale.
- Optimize MECM infrastructure for performance, reliability, and high availability.
- Create and manage MECM device collections, task sequences, applications, packages, and operating system deployment images.
- Automate MECM administrative tasks using PowerShell, APIs, and infrastructure-as-code tooling.
- Support Windows endpoint management, including patching strategies, compliance baselines, configuration items, and reporting.
- Improve endpoint reliability through monitoring, observability tooling, and proactive remediation.
- Integrate MECM with cloud-based services such as Intune, Azure AD, and Windows Update for Business.
- Troubleshoot complex MECM client and server issues using logs, diagnostics, and platform health data.
- Support enterprise-wide operating system deployments and in-place upgrades using MECM automation pipelines.
- Maintain MECM databases with relevant SQL Server administration and performance tuning.
- Contribute to system reliability through continuous improvement, infrastructure monitoring, and configuration drift reduction.
- Apply networking concepts supporting MECM components such as DP distribution, boundary groups, and PXE services.
- Work in a collaborative, forward-thinking, and innovation-driven environment.
- Apply Agile and DevSecOps/SRE concepts and best practices.
- Use Atlassian products (Jira, Confluence, Bitbucket) and create JIRA/Azure DevOps workflows, projects, and custom configurations.
- Administer/maintain SRE platforms via Ansible playbooks (e.g., upgrading Jenkins).
- Automate tasks with scripting languages like PowerShell or Python.
- Integrate and maintain third-party CI/CD tools like Jenkins and GitLab.
- Work with commercial cloud infrastructure deployment environments such as AWS and Azure.
- Use automated provisioning and configuration tools like Terraform, CloudFormation, Chef, Puppet, Ansible, or similar technologies.
- Apply knowledge of the Risk Management Framework (RMF) and DISA STIGs.
Requirements
- Active Secret clearance and ability to maintain it.
- BS degree and 5-10 years of prior relevant experience, or Master’s with 4-8 years of prior relevant experience.
- DoD 8570.01 IAT Level II Certification required prior to onboarding and must maintain it while supporting the contract.
- Ability to support program execution in classified environments and access SIPRNet from an Agency location on short notice (local travel).
- Experience with automated script design, coding, debugging, and maintenance (using bash, Python, etc.).
- Experience designing, deploying, and maintaining Microsoft Endpoint Configuration Manager (MECM) environments at enterprise scale.
- Proven ability to optimize MECM infrastructure for performance, reliability, and high availability.
- Hands-on experience creating and managing MECM device collections, task sequences, applications, packages, and operating system deployment images.
- Expertise in automating MECM administrative tasks using PowerShell, APIs, and infrastructure-as-code tooling.
- Strong background in Windows endpoint management, including patching strategies, compliance baselines, configuration items, and reporting.
- Demonstrated ability to improve endpoint reliability through monitoring, observability tooling, and proactive remediation.
- Experience integrating MECM with cloud-based services such as Intune, Azure AD, and Windows Update for Business.
- Ability to troubleshoot complex MECM client and server issues using logs, diagnostics, and platform health data.
- Experience supporting enterprise-wide operating system deployments and in-place upgrades using MECM automation pipelines.
- Knowledge of SQL Server administration relevant to maintaining MECM databases and enabling performance tuning.
- Experience contributing to system reliability through continuous improvement, infrastructure monitoring, and configuration drift reduction.
- Strong understanding of networking concepts supporting MECM components such as DP distribution, boundary groups, and PXE services.
- Knowledge of Agile and DevSecOps/SRE concepts and best practices.
- Hands-on experience with Atlassian products (Jira, Confluence, Bitbucket, etc.).
- Experience creating JIRA and/or Azure DevOps workflows, projects, and custom configurations.
- Experience administrating/maintaining SRE platforms via Ansible playbooks (e.g., upgrading Jenkins).
- Experience automating tasks with scripting languages like PowerShell or Python.
- Experience integrating/maintaining third-party CI/CD tools like Jenkins and GitLab.
- Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
- Experience with automated provisioning and configuration tools like Terraform, CloudFormation, Chef, Puppet, Ansible, or similar technologies.
- Working knowledge of the Risk Management Framework (RMF) and DISA STIGs.
Preferred Qualifications
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation for automating test environments.
- ITILv4, Scrum Master, or Agile SAFe certification(s) or applicable experience.
Pay
Salary Range: $85,000 - $110,000 per year. Compensation is based on responsibilities, education, experience, knowledge, skills, certifications, and other requirements.
Benefits
- 100% premium coverage for Medical, Dental, Vision, and Life Insurance.
- Supplemental Insurance.
- 401K matching.
- Flexible Time Off (PTO/Holidays).
- Higher Education/Training Reimbursement.
Schedule
Full-time, Hybrid.