Engineering Tech Lead (vNode)
About the role
At vCluster Labs, you aren't just shipping container runtime features; you are defining how Kubernetes operators get VM-grade tenant isolation without the VM tax. vNode replaces virtual kubelets and microVMs with a runtime built on Linux user namespaces and seccomp, and the person in this seat owns where that runtime goes next. You will partner directly with the vNode founding engineers, run the technical bar for the team, and ship the work that decides whether AI Clouds and regulated enterprises can adopt vNode as their default isolation layer.
Responsibilities
- Own the vNode technical execution: Drive the architecture for how vNode wraps containerd, integrates with the kubelet, and exposes safe isolation primitives. Set the bar for what ships, what gets deferred, and what gets redesigned.
- Go deep on container runtimes and isolation: Lead the work where vNode meets containerd, Kata Containers, gVisor, runc, and the kernel. Explain and improve exactly what happens between a Pod spec and a process running under a constrained user namespace with a tight seccomp profile.
- Ship the kubelet integration surface: Own how vNode plugs into the node lifecycle: CRI, kubelet device plugins, cgroups v2, eviction, and the rough edges between Kubernetes' node model and a runtime that does not assume one tenant per node.
- Raise the engineering bar: Run technical design reviews, set the pattern for testing isolation guarantees, and mentor the engineers shipping alongside you.
- Be Customer Zero for vNode: Run vNode against vCluster Platform tenant clusters internally before customers see it. Close the loop between what AI Cloud operators need and what vNode actually does in production.
- Represent vNode externally: Contribute upstream where it matters (containerd, runc, Kubernetes SIG-Node), write technical posts explaining why namespace-based isolation is the right answer, and represent vCluster Labs at KubeCon-class venues.
Requirements
- Deep container runtime experience: You have shipped production work against containerd directly, not just consumed it through Docker or Kubernetes. Direct experience with Kata Containers, gVisor, or another sandboxed/isolated runtime is a strong plus.
- Kubernetes node-level depth: You have worked inside the kubelet, the CRI layer, or a node-resident agent. You know what cgroups v2, OCI hooks, and the kubelet's PLEG do and where they break.
- Go systems programming chops: You write production Go for systems-level code (syscalls, namespaces, file descriptors, process lifecycle), not just service handlers.
- Linux isolation fluency: User namespaces, seccomp-bpf, capabilities, and Landlock are not abstract concepts; you have shipped against them and can reason about their failure modes.
- Tech Lead instincts: You set technical direction by writing the design doc, prototyping the hard part, and then bringing the team along. You raise the bar without becoming the bottleneck.
Qualifications
- Upstream contribution history: Meaningful commits to containerd, runc, Kata, gVisor, Kubernetes SIG-Node, or related projects.
- Tenant Isolation domain expertise: You have built or operated infrastructure where the threat model includes hostile workloads on shared hosts (AI Cloud operators, multi-tenant SaaS, regulated industries).
- Public technical voice: Talks, posts, or RFCs that move the conversation on container isolation.
About VCluster Labs
We are a venture-backed tech startup and the company pioneering Kubernetes virtualization for the AI era. We raised over $30M from top-tier VCs such as Khosla Ventures (first investor in OpenAI, GitLab, Stripe, Doordash) and are in a hyper-growth phase. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed globally with a remote-first work culture.
We are the leading platform for operating GPU infrastructure, enabling AI Cloud providers to deliver a hyperscaler-like experience to their customers and AI factories that need to build that same experience for their internal teams. Our platform delivers the full operational stack operators need to run their GPU data centers — managed Kubernetes, fast isolated tenant provisioning, and automated node provisioning and lifecycle management — enabling them to accelerate time to value, reduce operational burden, and maximize the ROI of every GPU.
We're the company behind vCluster, an open-source technology for virtualizing Kubernetes (10k+ GitHub stars, 40M+ virtual clusters created since 2021). Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI — a Kubernetes-native framework purpose-built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.
Benefits
- Competitive Salary: Competitive compensation package, including equity.
- Platinum-Level Insurance: Health, dental, vision, and life insurance, including plans for you and eligible dependents (benefits vary depending on country).
- Flexible Working Schedule: Results matter more than clocking in and out at the same time every day.
- Workplace Flexibility: Flexibility about where you work, with adjustments for life changes.
Culture & Values
- Make it Happen: Relentless bias for action and the grit to push through obstacles. Ruthlessly prioritize actions that drive measurable impact for the business.
- Own the Outcome: Responsibility doesn't end when a task is checked off; it ends when the value is delivered. Connect daily actions to the broader success of the company and customers.
- Create Wow: Measure success by the experience we generate, both inside and outside the company. Go the extra mile to support one another and drive each other to new heights.
- Open Source, Open Mind: Actively contribute to and maintain open-source projects. Foster meritocracy — the strongest ideas win, no matter who or where they come from.
- Build Tomorrow’s Standards, Intentionally: Define the state-of-the-art of tomorrow. Fearless in tearing down old approaches to build something better, but disciplined because users rely on our technology to run mission-critical infrastructure.
Pay
Compensation Range: $138,000 - $200,000