Machine Learning Engineer
About the role
UST is searching for a Machine Learning Engineer with experience in agent orchestration, tool/function calling, memory planning, and agent-to-agent communication.
The opportunity involves developing a multi-agent workload at the application level and understanding, profiling, and optimizing its execution—from the inference runtime and OMIX/middleware down to the compute runtime, Linux GPU driver, and accelerator hardware. You will optimize for latency, throughput, tokens/sec, memory footprint, and accelerator utilization, and select and tune models/inference engines based on agent workload characteristics, hardware capabilities, and performance requirements. Practical use of GitHub Copilot for development, debugging, and code generation is expected.
Responsibilities
- Hands-on development of multi-agent workloads
- Optimize multi-agent workload execution from application level down to accelerator hardware
- Select and tune models/inference engines based on agent workload characteristics and hardware capabilities
- Benchmark and profile AI workloads for latency, throughput, tokens/sec, GPU utilization, and memory bandwidth
- Identify performance bottlenecks across agent, model, inference engine, runtime, or driver layers
- Use GitHub Copilot for development, debugging, and code generation
Requirements
- Experience with agent frameworks such as LangGraph, LangChain, AutoGen, or CrewAI
- Strong understanding of LLMs, SLMs, and multimodal models, including Transformer architecture, attention, tokenization, context windows, and KV cache
- Hands-on experience with model selection, evaluation, and deployment for agentic workloads
- Understanding of model formats and optimization (ONNX, OpenVINO IR, safe tensors, quantization such as FP16/BF16/INT8/INT4)
- Experience with inference engines/runtime frameworks (OpenVINO, ONNX Runtime, vLLM, llama.cpp, TGI)
- Understanding of prefill vs. decode, batching, continuous batching, speculative decoding, KV-cache management, and memory optimization
- Understanding of CPU/GPU model execution, device placement, and heterogeneous inference
- Familiarity with GPU memory, kernel execution, synchronization, device selection, and host/device data movement
- Familiarity with middleware/accelerator compute runtimes (OMIX/OneAPI/SYCL, Level Zero, OpenCL)
- Strong Python skills and working knowledge of C/C++
- Proficiency with Git/GitHub, Linux shell, and debugging tools
- Docker/container fundamentals
Pay
Role Location: Oregon
Compensation Range: $82,000–$123,000
Benefits
- Full-time, regular employees accrue a minimum of 10 days of paid vacation per year, 6 days of paid sick leave, and 10 paid holidays
- Eligibility for paid bereavement leave and jury duty
- 401(k) Retirement Plan with employer matching
- Medical, dental, and vision insurance for employees and dependents residing in the US
- Company-paid Employee Only benefits: basic life insurance, accidental death and disability insurance, short- and long-term disability benefits
- Optional voluntary short-term disability benefits and participation in Health Savings Account (HSA), Flexible Spending Account (FSA) for healthcare, dependent child care, and/or commuting expenses
- Part-time employees receive 6 days of paid sick leave and are eligible for 401(k) with employer matching
- Full-time temporary employees receive 6 days of paid sick leave, 401(k) with employer matching, and medical, dental, and vision insurance
- Part-time temporary employees receive 6 days of paid sick leave
- All US employees receive paid sick leave benefits according to state or local laws where applicable