Machine Learning Engineer, Infrastructure
Bright Vision Technologies is a leading technology consulting and software development firm specializing in delivering innovative cloud, artificial intelligence, data, and enterprise solutions across the United States. The company fosters a collaborative, inclusive, and forward-thinking work environment that encourages continuous learning and professional growth, offering a dynamic workplace where professionals contribute to impactful projects shaping the future of AI and cloud computing.
About the role
We are seeking a highly skilled Machine Learning Infrastructure Engineer to join our remote team. In this role, you will design, build, and maintain high-performance inference platforms capable of supporting large-scale machine learning models in production environments. Your work will ensure the deployment of reliable, scalable, and efficient AI services, focusing on systems engineering aspects such as request routing, batching, caching, autoscaling, and GPU utilization. You will collaborate with ML researchers, data scientists, and product teams to optimize inference workflows, enhance system observability, and implement security measures.
Responsibilities
- Design, develop, and operate model serving platforms supporting diverse workloads such as LLMs, vision models, and recommendation systems
- Optimize inference performance through techniques like continuous batching, paged attention, speculative decoding, and request multiplexing
- Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints
- Build autoscaling and capacity management systems to balance latency, throughput, and operational costs
- Tune GPU utilization, manage memory hierarchies, and optimize KV cache strategies for large model inference
- Integrate model serving infrastructure with API gateways, identity management, and observability platforms
- Develop caching mechanisms, prompt deduplication, and response reuse strategies to enhance efficiency
- Establish comprehensive observability including latency metrics, queue dynamics, GPU utilization, and error tracking
- Create deployment workflows with canary releases, shadow testing, and automated rollback procedures
- Manage incident response for high-availability AI services and implement reliability improvements
- Collaborate with ML and product teams to support new model releases and feature rollouts
- Implement security controls such as request signing, content filtering, and abuse detection at the serving layer
- Document operational procedures, performance tuning guides, and system characteristics for internal teams
- Stay updated with the latest research in AI model serving and translate advances into production solutions
Qualifications
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
- Six or more years of experience in distributed systems, infrastructure, or ML platform engineering
- Proficiency in Python and systems programming languages such as Go, Rust, or C++
- Hands-on experience operating high-throughput, low-latency services in production environments
- Experience with large model inference frameworks like vLLM, TensorRT-LLM, or similar
- Deep understanding of GPU architectures, memory hierarchies, and accelerator utilization
- Familiarity with Kubernetes, cloud platforms, and autoscaling techniques
- Experience with observability tools including metrics collection, tracing, and structured logging
- Strong performance engineering and capacity planning skills
- Excellent communication skills and ability to respond effectively to incidents
Benefits
- Competitive salary range of $100,000 to $150,000 annually
- Remote work flexibility within the U.S.
- Opportunities for professional growth and career advancement
- Collaborative and innovative work environment
- Access to cutting-edge AI and cloud technologies
- Comprehensive health and wellness benefits
- Paid time off and leave policies
- Continuous learning and development programs