Cloud Engineer Senior
Position Overview
You are responsible for engineering, automating, and securing platform tools and secrets management infrastructure across enterprise environments. This role focuses on operational excellence for middleware platforms (Kafka, AMQ, Apigee, Control-M), secrets lifecycle management (HashiCorp Vault, AWS Secrets Manager), and driving AI-powered automation to modernize how platform services are provisioned, monitored, and governed.
Our Impact
We ensure the reliability, security, and operational efficiency of enterprise platform services that underpin hundreds of applications across Freddie Mac. By standardizing secrets management, automating platform operations, and leveraging AI for intelligent insights, we reduce operational risk, accelerate onboarding, and enable teams to focus on delivering business value rather than managing infrastructure.
Your Impact
- Automate platform tool operations (Confluent Kafka, Amazon MQ, Apigee, Control-M, Athena, Dynatrace) through Infrastructure as Code, custom tooling, and CI/CD pipeline integration.
- Design and implement AI-powered automation workflows using prompt engineering, LLM APIs, and agentic patterns to streamline platform operations, anomaly detection, and cost analysis.
- Build observability and operational dashboards for platform health, cost allocation, and adoption metrics across the tool portfolio.
- Standardize onboarding patterns for teams consuming platform services — secrets provisioning, connectivity.
- Support cloud migration efforts by ensuring platform tools are properly configured, monitored, and optimized in AWS environments.
Qualifications
- 5-7 years of experience in cloud infrastructure, platform engineering, or DevOps/SRE roles with production responsibilities; bachelor’s degree preferred.
- 3–5 years hands-on experience with secrets management platforms (HashiCorp Vault, AWS Secrets Manager, CyberArk) including policy authoring, dynamic secrets, and Kubernetes integration.
- Strong experience with AWS services in production (EKS, IAM/IRSA, S3, Secrets Manager, CloudWatch, Lambda) and Infrastructure as Code (Terraform, CloudFormation, Helm).
- Hands-on experience with at least 3 of the following: Confluent Kafka, Amazon MQ/ActiveMQ, Apigee API Gateway, Control-M scheduling, Dynatrace/APM.
- Proficiency in Python scripting and API automation — building integrations, data pipelines, and operational tooling.
- Demonstrated experience with AI/ML tools and prompt engineering — using LLM APIs (OpenAI-compatible, Claude, Bedrock) for automation, code generation, and operational intelligence.
- Experience with Kubernetes (EKS), service mesh (Istio), and container-based deployment patterns.
- Familiarity with CI/CD pipelines (Jenkins, Spinnaker, Gradle) and GitOps practices.
Preferred Qualifications
- Experience operating event-driven platform at scale – topic/queue provisioning, schema governance, consumer group monitoring, and capacity planning.
- Proven ability to build API-driven automation against platform management for inventory, compliance, and self-service workflows.
- Experience with agentic AI workflows, RAG (Retrieval-Augmented Generation), and tool-calling patterns.
- AWS certifications (Solutions Architect, DevOps Engineer, or Security Specialty).
Keys to Success in this Role
- Automate everything: treat manual processes as technical debt and build self-service patterns that scale.
- Stay hands-on: troubleshoot production issues, build prototypes, and iterate rapidly with working code.
- Leverage AI as a force multiplier: use prompt engineering and LLM integration to accelerate documentation, analysis, and automation development.
- Drive measurable outcomes with clear operational metrics (uptime, mean-time-to-provision, self-healing, cost optimization savings).
- Partner effectively across security, application, and infrastructure teams to ensure smooth adoption of platform services and governance standards.