Jobs · Engineering

AI Inference Tools Engineer

Modular · United States · 2 wk ago
RemoteRemoteEngineering$167k–$242k/yrFull-time

At Modular, we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges.

About the role

At Modular we are building a next generation AI platform to power modern applications and facilitate access to cutting-edge hardware. The MAX framework is our developer-facing layer: it defines the APIs developers use to express models, integrate custom kernels, orchestrate execution, iterate on quality and ship systems into production. As an engineer on the MAX Tools team you will build the developer tooling for Modular's AI inference stack to make the engineers who build and deploy AI models more productive. You will be working on tools to help people create, debug, and optimize models and kernels. Modular’s vertically integrated ML serving stack allows you to solve a uniquely diverse set of problems, and gives you especially broad scope for creating solutions. A growing part of our work is agent-native tooling: tools designed so that AI agents can profile, triage, and debug issues under engineer guidance.

Candidates based in the US, Canada, or UK are welcome to apply. To support growth and collaboration, those in earlier career stages work in a hybrid capacity at our Los Altos, CA or Edinburgh, UK office (minimum 3 days per week on-site) with relocation assistance provided for out-of-state candidates based in the US. Senior members have both in office or remote flexibility. Onboarding for new hires is conducted in person at an appropriate office.

Responsibilities

  • Enhance the MAX developer experience by creating tools that can be part of traditional and agentic workflows, by surfacing data through the entire AI stack, and by supporting a widening list of heterogeneous hardware platforms.
  • Design, build, and maintain key technologies that support the MAX framework, such as profiling and tracing tools, debuggers and associated tools to check correctness, and metrics and dashboards to monitor the performance and health of MAX deployments.
  • Create agent-native tooling: tools that can be used by humans and agents alike, and the agentic workflows that go around them. Write agents to bring up models, to triage issues with running systems, and to profile and optimize based on real workloads.
  • Connect MAX tools through the entire stack: by strengthening the tooling interfaces between MAX, Mojo, and the driver APIs they run on; and by integrating tool use cleanly with the orchestration layers above.
  • Participate in design discussions and code reviews to uphold high engineering standards.
  • Work directly with the kernel, model, and serving engineers who use your tools every day, and turn their recurring pain points into better tools.

Requirements

  • Passion for creating exceptional developer experiences and an understanding of what makes great tools feel effortless.
  • Experience designing, building, and shipping developer tools, developer infrastructure, or observability systems that other engineers rely on.
  • Fluency in one or more systems/performance languages (C++, Rust, Go) and one or more user-facing languages (Python; familiarity with Mojo is a plus).
  • Experience measuring software performance via profiling, tracing, or benchmarking, and the systems knowledge to analyze those results.
  • Clear written communication and the ability to work independently in a distributed team.
  • A user-focused mindset: you seek out the people who use your tools, learn their workflows, and prioritize what makes them more productive.

Skills

  • Experience with GPU performance tooling (Nsight Systems/Compute, rocprof, Perfetto, or similar) or GPU programming (CUDA, ROCm).
  • Knowledge of basic AI implementation and modeling techniques and familiarity with AI frameworks like PyTorch, JAX, or TensorFlow.
  • Experience building tooling for AI agents (e.g. agent skills, MCP).

Benefits

  • Premier insurance plans, up to 5% 401k matching, flexible paid time off, and more (specific benefit packages may vary based on location).
  • Competitive compensation including stock options.
  • Team building events, including regular team onsites and local meetups in Los Altos, CA and other cities (traveling 2-4 times a year is expected).

Pay

The estimated base salary range for this role to be performed in the US, regardless of the state, is $167,000.00 - $242,000.00 USD. The estimated base salary range for this role to be performed in the UK is £97,000 - £140,000 GBP. The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, including but not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. The total compensation for a candidate will also include annual target bonus, equity, and benefits, with equity making up a significant portion of your total compensation.

Similar jobs