DevOps Engineer
Skiffra · Los Angeles, CA · 3 days ago
RemoteRemoteEngineeringContract
About Skiffra
Skiffra is building the intelligence layer for the physical world. We translate complex, real-world environments into clear, actionable data. While most AI companies focus on digital industries, we design AI native orchestration systems for the sectors that extract, move, and build ecosystems around us. Our first proving ground is mining and natural resources, but the platform is modular by design and built to scale into any industry where decisions are expensive and the consequences are real.
What You'll Do
- Own our AWS infrastructure and cloud architecture, ensuring it's secure, scalable, and highly available
- Build and maintain Kubernetes clusters and the tooling required to deploy and operate production services
- Create internal developer tooling that makes it easy for engineers to build, test, and deploy safely
- Design and improve CI/CD pipelines that enable rapid, reliable releases
- Partner with backend, data, and AI engineers to build infrastructure that supports distributed systems and production AI workloads
- Own observability across the platform, including monitoring, logging, alerting, and incident response
- Continuously improve security, reliability, performance, and cost efficiency across our infrastructure
- Help establish engineering best practices around infrastructure, deployments, and operational excellence
- Play a key role in shaping both our technical roadmap and engineering culture as one of our earliest hires
We're Looking For
- 8+ years of experience building and operating cloud infrastructure in production environments
- Strong experience with AWS, Kubernetes, and Terraform
- Excellent Python skills for automation, tooling, and infrastructure development
- Experience managing PostgreSQL and cloud-native databases
- A strong understanding of containers, networking, Linux, security, and distributed systems
- Experience designing scalable, highly available production infrastructure
- Comfortable debugging complex production issues and owning systems from design through operation
- Excellent communication skills and the ability to collaborate across engineering disciplines
- Experience at an early-stage startup where you've built infrastructure from the ground up is highly preferred
- Comfortable operating across AWS and Azure (GCP a plus); experience with per-client / multi-cloud deployment topologies
- Experience operating data pipelines and lakehouse infra, object-store table formats (Iceberg), a catalog (Glue/Nessie), an orchestrator (Dagster/Airflow/Temporal), dbt in CI, and query engines (DuckDB/Trino)
- MLOps/LLMOps, GPU infrastructure, inference serving, LLM observability/evals, cost controls for AI workloads
- Multi-tenant isolation, data-residency-aware deployments, and SOC 2 / ISO 27001 / GDPR experience
Nice to Have
- Experience with observability tooling such as Prometheus, Grafana, OpenTelemetry, Datadog, or CloudWatch
- Familiarity with DynamoDB, event-driven architectures, and distributed messaging systems
- Experience improving developer experience through platform engineering and internal tooling
- Previous experience as a founding engineer or technical lead