Site Reliability Engineer
BAM Ventures · New York, NY · 4 days ago
EngineeringFull-time
About the Company
Aisle builds a real-time control layer on top of physical retail, connecting shopper actions, incentives, and outcomes into a continuous feedback loop. This enables brands to trigger, measure, and optimize retail performance in real time—laying the foundation for autonomous retail execution. Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar.
Under the hood, Aisle handles high-volume event ingestion, real-time decisioning, and low-latency infrastructure to influence behavior in the moment. The team is scaling the system to manage increased volume, complexity, and automation, including AI systems integrated into the production path.
About You
You’ll thrive at Aisle if you:
- Enjoy communicating across technical and non-technical teams to build shared understanding and collaborative action
- Solve problems in stages—balancing immediate triage, temporary patches, and long-term fixes to increase system durability
- Are comfortable moving quickly, shipping improvements, and iterating in production
- Treat AI as a core part of your workflow, using tools beyond basic copilots to design, debug, and ship systems faster
- Approach AI agents as chaotic, untrusted components, prioritizing isolation, rate limiting, and permissions scoping
- Have experience building or orchestrating AI/agent workflows in production or through serious side projects
Responsibilities
Reliability & Infrastructure
- Own the reliability, scalability, and observability of infrastructure across GCP and Vercel
- Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards
- Manage IAM, service accounts, and security best practices across the cloud environment
- Participate in on-call rotation, incident response, and post-mortems, turning learnings into systemic improvements
Core Infrastructure & State
- Stabilize core infrastructure and stateful systems under heavy concurrent load from consumer traffic and agent-driven workflows
- Build and maintain event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments
- Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution
- Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)
- Automate infrastructure provisioning, deployments, and operational workflows
AI & Next-Gen Tooling: Agent Ops
- Build agent operations infrastructure to enable AI agents to run safely and reliably in production
- Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry
- Define CI/CD for an agent-driven engineering environment, including code shipping, review, validation, and deployment with agents in the loop
- Own visibility into AI usage, reliability, and spend as the agent footprint scales
Cross-functional Impact
- Partner closely with engineering and product teams to maintain reliability without slowing development velocity
- Act as a force multiplier across the team, helping engineers ship faster and more safely
Requirements
Must Haves
- 4+ years in SRE, DevOps, or infrastructure/platform engineering
- Strong, hands-on experience with a major cloud platform (preferably GCP)
- Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)
- Solid understanding of IAM, security, and cloud best practices
- Experience with observability tools like Datadog
- Familiarity with Node.js environments
- AI is part of your daily engineering workflow—you use AI tools thoughtfully to improve velocity, reliability, and operational excellence beyond basic code generation
- Experience building or orchestrating AI/agent workflows (work or serious personal projects)
- High ownership, strong curiosity, and a bias toward action
Nice to Haves
- Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls
- Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns
- Hands-on experience with GCP Workflows for orchestration
- Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling
- Experience with Vercel deployment and edge computing
- Familiarity with the Kubernetes ecosystem
- Familiarity with Redis and BullMQ
- Understanding of SOC 2 compliance requirements and implementation
- Previous experience in a high-growth startup environment
- Previous backend engineering experience to bridge the gap between infrastructure and code
About the Stack
- Cloud: Google Cloud Platform (GCP)
- Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes
- Database: PostgreSQL with pgbouncer, Prisma
- Observability: Datadog
- Runtime: Node.js, TypeScript
- Deployment: Vercel, GCP
- Frontend: React, Next.js, TypeScript
- Backend: TypeScript, PostgreSQL, Next.js