Platform Engineer IV
Eliassen Group · Greenwood Village, CO · 3 days ago
Information Technology$70–$80/hrContract
Hybrid 2 onsite / 3 WFH in Greenwood Village, CO.
About the role
Design, build, and maintain the AWS infrastructure that underpins the IIA Data Lake, agent runtime environments, CI/CD pipelines, and graph database systems. Develop the application code, automation, and internal tooling for those systems. Ensure production environments are stable, scalable, and secure while enabling data science and agentic AI workloads to operate reliably at scale.
Responsibilities
- Design and manage AWS infrastructure for the IIA Data Lake including S3, Glue, Athena, and EMR.
- Manage cross-account connectivity, VPC networking, security groups, and IAM roles and policies.
- Build and maintain infrastructure for AI agent runtime environments for LangGraph agents via LangSmith Deployments.
- Design and implement event-driven infrastructure with triggers, queues, and pub or sub to orchestrate agents and data workflows.
- Support deployment and operation of AWS Neptune for a network topology graph including capacity and performance tuning.
- Implement infrastructure as code using Terraform or CloudFormation for repeatable provisioning.
- Manage IAM key rotations, secrets management, and security compliance for on-premises and cloud integrations.
- Build and maintain CI or CD pipelines for agent deployments using GitLab CI or CD, Docker, and Artifactory.
- Manage container lifecycle including image builds, versioning, and rollbacks for LangSmith Deployments.
- Automate deployment workflows to promote agents from development to production.
- Coordinate with platform teams on AI gateway integration, cross-account deployment, and connectivity.
- Support and maintain ETL pipelines in Scala or Spark, and onboard new data sources.
- Develop Python scripts, CLIs, services, and automation utilities for ingestion, deployment, and operations.
- Write integration code and glue services for upstream and downstream systems and external platforms.
- Ensure production stability with monitoring, alerting, incident response, and SLA adherence.
- Implement monitoring and alerting for deployed agents including health, errors, latency, and resource use.
- Coordinate connectivity and firewall requests with upstream data teams and platform teams.
- Support infrastructure needs for new data onboarding and scheduler maintenance with Airflow and event triggers.
- Perform other duties as required.
Requirements
- Expert experience with AWS services including EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, and CloudWatch.
- Strong infrastructure as code experience with Terraform or CloudFormation.
- Background in cross-account AWS architectures, VPC peering, PrivateLink, and transit gateway.
- IAM policy design with least privilege and service account management.
- Containerization with Docker and container orchestration.
- CI or CD pipelines with GitLab CI.
- Event-driven architectures and messaging such as Kafka, SNS or SQS, or EventBridge.
- Linux proficiency and shell scripting.
- Monitoring and alerting with CloudWatch, Prometheus, Grafana, or similar.
- Networking fundamentals including DNS, CIDR, NAT, VPC endpoints, firewalls, and security groups.
- Cross-team coordination on connectivity and access requirements.
- Proficiency in Python or a comparable language for automation and tooling.
- Software engineering fundamentals including Git workflows, code review, modular design, dependency management, and documentation.
- Automated tests for application and infrastructure code and CI or CD integration.
- Integration against REST APIs and cloud SDKs such as boto3.
Preferred Qualifications
- Graph databases such as AWS Neptune or Neo4j.
- Apache Kafka or similar streaming platforms.
- Apache Spark with Scala for distributed and streaming processing.
- Airflow or similar orchestration platforms.
- Telecommunications or large-scale network operations experience.
- Familiarity with AI or ML infrastructure such as model serving, compute, and artifact management.
- Splunk integration including Edge Processor and MCP connectivity.
- AWS certifications such as Solutions Architect or DevOps Engineer.
- Experience developing and operating small services or APIs with FastAPI or Flask.
Education Requirements
- Bachelor's degree in Computer Science, Information Technology, Systems Engineering, or related field, or relevant experience.
- Bachelor's degree: 5+ years of platform or infrastructure engineering experience.
- Master's degree: 3+ years of platform or infrastructure engineering experience.
Benefits
- Medical, Dental, and Vision benefits.
- 401k with company matching.
- Life insurance.
Pay
$70.00 to $80.00/hr. W2.