Senior Research Scientist- Vision-Language-Action (VLA) Models
Find Data Science Jobs · Sunnyvale, CA · 6 days ago
OTHR$185k–$215k/yrFull-time
About the role
As a Senior Research Scientist specializing in Vision-Language-Action (VLA) Models, you will contribute to research projects at the forefront of the ADAS/AD industry, advancing Embodied AI for domains such as ADAS/AD, industrial automation, and robotics.
Responsibilities
- Conduct research and engineering in core AI and machine learning fields, including computer vision, autonomous planning, and open-world learning.
- Push the boundaries in modular end-to-end perception and planning for ADAS/AD, incorporating advancements in large vision-language-(action) models to enhance reasoning and explainability.
- Collaborate cross-functionally with global research and engineering teams to ensure seamless technology transfer and system integration.
- Implement research results to solve real-world challenges, integrating them into Bosch’s existing platforms.
- Stay at the forefront of innovation by engaging with academic and industry communities through conferences, workshops, and technical events.
- Document and disseminate research findings through high-caliber publications and/or patent submissions.
Qualifications
Basic Qualifications
- Ph.D. in Computer Science, Robotics, or a related discipline, or a Master’s degree with 2–4 years of industry experience post-graduation.
- A minimum of 5 years of R&D experience (or equivalent graduate research) in AI technologies, including Computer Vision and Robotic or Automotive Motion and Behavioral Planning.
- Proficiency in programming languages commonly used in machine learning (e.g., Python, C++, Rust).
- Strong interpersonal, communication, and teamwork capabilities.
- Knowledge of major machine learning frameworks like TensorFlow or PyTorch.
- Hands-on experience in reinforcement learning for behavior or motion planning, with familiarity in techniques such as PPO, DQN, or DDPG.
- A strong portfolio of publications in premier machine learning, deep learning, robotics, and computer vision journals and conferences.
Preferred Qualifications
- Experience with real-world product development and deployment of autonomous systems.
- Hands-on experience building and applying multimodal transformer-based sequence-to-sequence models, particularly vision-language-action models.
- Expertise in computer vision and deep learning, with work in areas such as multimodal transformers, multimodal language models, diffusion models, NeRF, Gaussian splatting, object detection/segmentation, 3D scene understanding, sensor calibration, SfM, or voxel/BEV grid-based feature representation.
Pay
Competitive base salary range in US-California: $185,000 – $215,000, with additional annual corporate bonus and long-term incentive bonus based on sustained impact and contribution.
Benefits
- Premium health coverage.
- 401(k) with generous employer matching.
- Resources for financial planning and goal setting.
- Ample paid time off and parental leave.
- Comprehensive life and disability protection.