AI Platform Engineer
Bright Vision Technologies · Chapel Hill, NC · Yesterday
RemoteRemoteEngineering$100k–$150k/yrFull-time
About the Role
Bright Vision Technologies is seeking a highly experienced AI Platform Engineer to design, build, and operate enterprise-scale AI inference and machine learning platforms. This 100% remote, full-time position offers a salary range of $130,000–$180,000 annually, based on experience. U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Responsibilities
- Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments
- Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services li>Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads
- Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services
- Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices
- Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance
- Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures
- Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers to deploy and support production AI models
- Drive cloud infrastructure optimization, resource utilization, FinOps initiatives, and operational excellence
- Mentor engineering teams, conduct architecture reviews, and establish best practices for AI platform engineering and cloud-native development
- Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to drive continuous innovation
Qualifications
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline
- 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering
- Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++
- Extensive experience with Large Language Model (LLM) serving, model inference optimization, and production AI infrastructure
- Hands-on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks
- Strong expertise in Kubernetes, container orchestration, Docker, and cloud-native application architectures
- Experience optimizing GPU workloads using CUDA, NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure
- Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP)
- Strong understanding of distributed systems, networking, scalability, observability, and security best practices
- Excellent analytical, communication, collaboration, and technical leadership skills
Preferred Qualifications
- Experience designing and operating multi-region AI platforms and globally distributed inference services
- Knowledge of model optimization techniques such as quantization, pruning, compression, speculative decoding, KV cache optimization, and mixed-precision inference
- Experience with MLOps, GitOps, Infrastructure as Code (Terraform, Bicep, CloudFormation), and CI/CD automation
- Familiarity with service mesh technologies such as Istio or Linkerd, API gateways, and event-driven architectures
- Contributions to open-source AI infrastructure projects, technical publications, patents, or conference presentations
- Experience implementing FinOps strategies, cloud cost optimization, and enterprise AI governance
- Experience with multi-region AI deployments and AI infrastructure
- Familiarity with model optimization techniques such as quantization or compression
- Open-source contributions or experience supporting large-scale AI APIs