it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

2 jobs found

Email me jobs like this
Refine Search
Current Search
senior machine learning engineer robotics
Staff Machine Learning Engineer - Vision-Language Foundation Models
Waymo Mountain View, California
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The Team & Mission: In the Oracle Perception team, our mission is to build the ultimate cognitive engine for autonomous driving. We are pioneering the use of large multimodal foundation models (e.g., Gemini) to build a powerful offboard reasoning and data flywheel system. We are moving beyond traditional perception to true scene understanding and driving actions-building offboard models that can comprehend complex driving problems, predict object/scene dynamics, and deduce driving paths with logical rationale. Our core focus is advancing the VLM foundation itself. By pushing the boundaries of multimodal pre-training and state-of-the-art post-training (SFT, RL) , we are creating models capable of rich, reasoning-based autolabeling at a massive scale. This closed-loop data engine directly powers the training and evolution of Waymo's real-time onboard models. If you are passionate about defining VLM training recipes, scaling laws, and unlocking complex reasoning via RL, this is your opportunity to redefine the foundation of autonomous driving. In this hybrid role, you will report to a Senior Staff Technical Lead Manager. You Will: Drive Pre-training & Domain Adaptation: Lead the technical strategy for curating and constructing massive-scale, high-quality multimodal pre-training datasets. Define data mixture strategies to instill deep, Waymo-specific driving intuition and physics-grounded understanding into foundation models without catastrophic forgetting. Lead Post-Training & Reasoning Enhancement: Design and implement state-of-the-art fine-tuning (SFT) and Reinforcement Learning (RLHF/RLAIF, DPO/GRPO/PPO) pipelines. Drastically improve the model's instruction-following and complex reasoning capabilities (e.g., Chain-of-Thought, spatial-temporal reasoning, and driving rationale prediction). Pioneer the VLM Data Flywheel: Architect the highly scalable inference and evaluation pipelines that leverage these trained Gemini-class models to autonomously source, sample, and autolabel critical edge cases, directly accelerating the onboard perception models. Define Training Recipes & Scaling Laws: Conduct rigorous ablation studies to optimize model architectures, token budgets, and loss functions. Establish best practices for scaling multimodal training efficiently on large GPU/TPU clusters. Drive Cross-Functional AI Strategy: Act as the principal technical visionary across ML Infra, Perception, Behavior, and AI Foundation teams. Drive consensus on the data flywheel architecture and embed VLM reasoning capabilities seamlessly into the broader autonomous vehicle stack. Provide Staff-Level Technical Leadership: Own the long-term technical roadmap for foundation model development. Mentor senior engineers, lead rigorous design reviews, and establish standard-setting engineering practices from advanced prototyping to production deployment. You Have: Master's degree in Computer Science, AI, ML, or a related technical field. 8+ years of hands-on experience designing, training, and scaling deep learning models, with at least 3+ years focused deeply on training Large Language Models (LLMs) or Vision-Language Models (VLMs) . Proven expertise in the full lifecycle of Foundation Models: from pre-training data curation (interleaved formats, tokenization) and distributed training to advanced post-training techniques. Expert-level understanding of training infrastructure and distributed paradigms (e.g., FSDP, Megatron, JAX/Pax) required for training massive models reliably. Expert-level software engineering fundamentals using Python, PyTorch, or JAX, with a track record of building reliable, highly scalable ML systems. Proven ability to operate with high ambiguity, define technical roadmaps, and drive complex, multi-quarter technical initiatives across multiple teams in a fast-paced environment. We Prefer: PhD in Computer Science, Artificial Intelligence, or a related field. Strong publication record in top-tier AI venues (e.g., NeurIPS, ICML, ICLR, CVPR) focusing on foundation models, large-scale training, reinforcement learning, or reasoning. Deep experience with advanced Reinforcement Learning paradigms applied to language or vision tasks ( focusing on improving System 2 thinking, logical deduction, and model alignment ). Demonstrated experience in Data Engineering for Foundation Models at the scale of billions/trillions of tokens (e.g., deduplication, quality filtering, synthetic data generation). Familiarity with the systemic challenges of multimodal perception in robotics or autonomous driving (e.g., 3D scene understanding, trajectory prediction). A proven track record of Staff-level impact: influencing product direction, pioneering zero-to-one ML architectures, and multiplying team efficiency through technical leadership. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $251,000-$310,000 USD
09/23/2026
Full time
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. The Team & Mission: In the Oracle Perception team, our mission is to build the ultimate cognitive engine for autonomous driving. We are pioneering the use of large multimodal foundation models (e.g., Gemini) to build a powerful offboard reasoning and data flywheel system. We are moving beyond traditional perception to true scene understanding and driving actions-building offboard models that can comprehend complex driving problems, predict object/scene dynamics, and deduce driving paths with logical rationale. Our core focus is advancing the VLM foundation itself. By pushing the boundaries of multimodal pre-training and state-of-the-art post-training (SFT, RL) , we are creating models capable of rich, reasoning-based autolabeling at a massive scale. This closed-loop data engine directly powers the training and evolution of Waymo's real-time onboard models. If you are passionate about defining VLM training recipes, scaling laws, and unlocking complex reasoning via RL, this is your opportunity to redefine the foundation of autonomous driving. In this hybrid role, you will report to a Senior Staff Technical Lead Manager. You Will: Drive Pre-training & Domain Adaptation: Lead the technical strategy for curating and constructing massive-scale, high-quality multimodal pre-training datasets. Define data mixture strategies to instill deep, Waymo-specific driving intuition and physics-grounded understanding into foundation models without catastrophic forgetting. Lead Post-Training & Reasoning Enhancement: Design and implement state-of-the-art fine-tuning (SFT) and Reinforcement Learning (RLHF/RLAIF, DPO/GRPO/PPO) pipelines. Drastically improve the model's instruction-following and complex reasoning capabilities (e.g., Chain-of-Thought, spatial-temporal reasoning, and driving rationale prediction). Pioneer the VLM Data Flywheel: Architect the highly scalable inference and evaluation pipelines that leverage these trained Gemini-class models to autonomously source, sample, and autolabel critical edge cases, directly accelerating the onboard perception models. Define Training Recipes & Scaling Laws: Conduct rigorous ablation studies to optimize model architectures, token budgets, and loss functions. Establish best practices for scaling multimodal training efficiently on large GPU/TPU clusters. Drive Cross-Functional AI Strategy: Act as the principal technical visionary across ML Infra, Perception, Behavior, and AI Foundation teams. Drive consensus on the data flywheel architecture and embed VLM reasoning capabilities seamlessly into the broader autonomous vehicle stack. Provide Staff-Level Technical Leadership: Own the long-term technical roadmap for foundation model development. Mentor senior engineers, lead rigorous design reviews, and establish standard-setting engineering practices from advanced prototyping to production deployment. You Have: Master's degree in Computer Science, AI, ML, or a related technical field. 8+ years of hands-on experience designing, training, and scaling deep learning models, with at least 3+ years focused deeply on training Large Language Models (LLMs) or Vision-Language Models (VLMs) . Proven expertise in the full lifecycle of Foundation Models: from pre-training data curation (interleaved formats, tokenization) and distributed training to advanced post-training techniques. Expert-level understanding of training infrastructure and distributed paradigms (e.g., FSDP, Megatron, JAX/Pax) required for training massive models reliably. Expert-level software engineering fundamentals using Python, PyTorch, or JAX, with a track record of building reliable, highly scalable ML systems. Proven ability to operate with high ambiguity, define technical roadmaps, and drive complex, multi-quarter technical initiatives across multiple teams in a fast-paced environment. We Prefer: PhD in Computer Science, Artificial Intelligence, or a related field. Strong publication record in top-tier AI venues (e.g., NeurIPS, ICML, ICLR, CVPR) focusing on foundation models, large-scale training, reinforcement learning, or reasoning. Deep experience with advanced Reinforcement Learning paradigms applied to language or vision tasks ( focusing on improving System 2 thinking, logical deduction, and model alignment ). Demonstrated experience in Data Engineering for Foundation Models at the scale of billions/trillions of tokens (e.g., deduplication, quality filtering, synthetic data generation). Familiarity with the systemic challenges of multimodal perception in robotics or autonomous driving (e.g., 3D scene understanding, trajectory prediction). A proven track record of Staff-level impact: influencing product direction, pioneering zero-to-one ML architectures, and multiplying team efficiency through technical leadership. The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $251,000-$310,000 USD
Staff Machine Learning Operations Engineer - Computer Vision
ATI Woburn, Massachusetts
Job Description Job Description About ATI: Automated Tire (ATI) is a Series-B startup revolutionizing automotive service with innovative robotic and software technology. Founded by experienced entrepreneurs and backed by major players in the automotive and tire sectors, ATI is building the next generation of tools that make tire shops and dealership service lanes faster, safer, and smarter. If you're passionate about building products that ship into real-world environments, ATI is the place for you. Position Overview: BrakeWise is our production brake inspection product: a mobile application paired with a camera probe that technicians use to assess pad and rotor condition during live service work. The machine learning behind it is a multi-stage pipeline of segmentation and classification models that turn raw imagery into a wear assessment a shop can act on and charge for. That pipeline works, and it is an MVP. It runs on Cloud Functions, and it will not carry us to the customer volume we're signing. We're looking for a Staff MLOps Engineer to own it - to take it from a working prototype to a serving architecture that holds up under real throughput, with the latency, cost, and reliability characteristics a paying customer expects. You'll own every aspect of how our models reach production and how they get better: serving infrastructure, deployment and rollback, monitoring and drift detection, the retraining loop, and the evaluation discipline that tells us whether a new model is actually an improvement. Model accuracy here has commercial consequences - a bad wear call is either a missed repair or an unnecessary one, in front of a customer. This is also the senior cloud architecture voice on the team. You'll partner closely with our Staff Full Stack Engineer, who owns the mobile app and customer dashboard, reviewing designs and setting GCP practices across the platform rather than only within the ML stack. Responsibilities: Own the multi-stage inference pipeline (segmentors and classifiers) end to end - serving architecture, latency, throughput, reliability, and cost per inspection Re-architect the pipeline off its current Cloud Functions MVP onto infrastructure that scales: containerized inference, GPU-backed or accelerated serving where it pays for itself, queueing, batching, and autoscaling Own model deployment: versioning, staged rollout, canary and shadow evaluation, and fast rollback when a model regresses Build and own the improvement loop - field data collection, labeling workflows, dataset versioning, evaluation harnesses, and regression suites that catch quality loss before customers do Monitor model quality in production: drift detection, segmented performance analysis, and triage of real-world failures against real inspection imagery Define the metrics that matter commercially - false-positive and false-negative rates on a wear call, technician override rate, unit inference cost - and report against them Improve model performance directly: architecture selection, augmentation, hard-example mining, and quantization or distillation where latency and cost demand it Evaluate on-device versus cloud inference trade-offs for the mobile app, and own whichever path we choose Establish MLOps foundations: reproducible training, experiment tracking, CI/CD for models, and infrastructure as code Serve as the cloud architecture counterpart to the Staff Full Stack Engineer - reviewing designs, setting GCP best practices, and raising the platform's infrastructure bar Work with hardware and field operations on capture quality - lighting, focus, and probe positioning - since upstream image quality sets the ceiling on model performance Own production support for the ML stack, including incident response and on-call participation for inference availability Proactively identify technical risks and architectural trade-offs, and communicate them clearly to leadership Requirements 8+ years of professional engineering experience, including several years owning machine learning systems in production - not solely model development Demonstrated experience taking a computer vision pipeline from prototype to production scale, serving real users at meaningful volume Deep experience deploying and operating segmentation and classification models, including multi-stage pipelines where one model's output feeds the next Strong cloud infrastructure background, preferably GCP - Vertex AI, Cloud Run, GKE, Cloud Functions, Cloud SQL, Pub/Sub, and Docker Production-grade Python, and fluency with PyTorch or TensorFlow Hands-on experience with model serving and optimization - Triton, TorchServe, ONNX, TensorRT, quantization, or equivalent Experience owning deployment and support for a live system, including incident response, rollback, and on-call Experience building data and labeling pipelines with dataset versioning and reproducible evaluation Comfort with infrastructure as code (e.g., Terraform) and CI/CD automation (e.g., GitHub Actions) Sound judgment on the accuracy, latency, and cost trade-offs that determine whether an ML product is viable Excellent problem-solving, debugging, and communication skills, including with non-technical stakeholders Preferred Qualifications: On-device or edge inference experience (Core ML, TensorFlow Lite, ExecuTorch) and integration into mobile applications Active learning or human-in-the-loop labeling systems Computer vision on small, long-tail, or industrial inspection datasets rather than large public benchmarks Experience with camera and sensor integration, or working alongside hardware teams on capture quality Experience with robotics, IoT, or edge computing (ROS or similar platforms) Familiarity with automotive service, dealership operations, or DMS ecosystems Contributions to open-source projects Why Join ATI: Be part of a groundbreaking startup transforming automotive service technology Work with a team of industry veterans and top-tier robotics and software talent Own the ML platform for a product that already has paying customers - your architecture decisions set the scaling ceiling Our customers are our investors, so you'll develop and test in real service lane environments A genuine data advantage: proprietary inspection imagery from real shops that no public dataset can replicate Clear Total Addressable Market with strong pull from B2B partners Competitive salary and comprehensive benefits package Prime location in Woburn, MA with on-site parking Collaborative, low-ego, high-intensity work environment
09/18/2026
Full time
Job Description Job Description About ATI: Automated Tire (ATI) is a Series-B startup revolutionizing automotive service with innovative robotic and software technology. Founded by experienced entrepreneurs and backed by major players in the automotive and tire sectors, ATI is building the next generation of tools that make tire shops and dealership service lanes faster, safer, and smarter. If you're passionate about building products that ship into real-world environments, ATI is the place for you. Position Overview: BrakeWise is our production brake inspection product: a mobile application paired with a camera probe that technicians use to assess pad and rotor condition during live service work. The machine learning behind it is a multi-stage pipeline of segmentation and classification models that turn raw imagery into a wear assessment a shop can act on and charge for. That pipeline works, and it is an MVP. It runs on Cloud Functions, and it will not carry us to the customer volume we're signing. We're looking for a Staff MLOps Engineer to own it - to take it from a working prototype to a serving architecture that holds up under real throughput, with the latency, cost, and reliability characteristics a paying customer expects. You'll own every aspect of how our models reach production and how they get better: serving infrastructure, deployment and rollback, monitoring and drift detection, the retraining loop, and the evaluation discipline that tells us whether a new model is actually an improvement. Model accuracy here has commercial consequences - a bad wear call is either a missed repair or an unnecessary one, in front of a customer. This is also the senior cloud architecture voice on the team. You'll partner closely with our Staff Full Stack Engineer, who owns the mobile app and customer dashboard, reviewing designs and setting GCP practices across the platform rather than only within the ML stack. Responsibilities: Own the multi-stage inference pipeline (segmentors and classifiers) end to end - serving architecture, latency, throughput, reliability, and cost per inspection Re-architect the pipeline off its current Cloud Functions MVP onto infrastructure that scales: containerized inference, GPU-backed or accelerated serving where it pays for itself, queueing, batching, and autoscaling Own model deployment: versioning, staged rollout, canary and shadow evaluation, and fast rollback when a model regresses Build and own the improvement loop - field data collection, labeling workflows, dataset versioning, evaluation harnesses, and regression suites that catch quality loss before customers do Monitor model quality in production: drift detection, segmented performance analysis, and triage of real-world failures against real inspection imagery Define the metrics that matter commercially - false-positive and false-negative rates on a wear call, technician override rate, unit inference cost - and report against them Improve model performance directly: architecture selection, augmentation, hard-example mining, and quantization or distillation where latency and cost demand it Evaluate on-device versus cloud inference trade-offs for the mobile app, and own whichever path we choose Establish MLOps foundations: reproducible training, experiment tracking, CI/CD for models, and infrastructure as code Serve as the cloud architecture counterpart to the Staff Full Stack Engineer - reviewing designs, setting GCP best practices, and raising the platform's infrastructure bar Work with hardware and field operations on capture quality - lighting, focus, and probe positioning - since upstream image quality sets the ceiling on model performance Own production support for the ML stack, including incident response and on-call participation for inference availability Proactively identify technical risks and architectural trade-offs, and communicate them clearly to leadership Requirements 8+ years of professional engineering experience, including several years owning machine learning systems in production - not solely model development Demonstrated experience taking a computer vision pipeline from prototype to production scale, serving real users at meaningful volume Deep experience deploying and operating segmentation and classification models, including multi-stage pipelines where one model's output feeds the next Strong cloud infrastructure background, preferably GCP - Vertex AI, Cloud Run, GKE, Cloud Functions, Cloud SQL, Pub/Sub, and Docker Production-grade Python, and fluency with PyTorch or TensorFlow Hands-on experience with model serving and optimization - Triton, TorchServe, ONNX, TensorRT, quantization, or equivalent Experience owning deployment and support for a live system, including incident response, rollback, and on-call Experience building data and labeling pipelines with dataset versioning and reproducible evaluation Comfort with infrastructure as code (e.g., Terraform) and CI/CD automation (e.g., GitHub Actions) Sound judgment on the accuracy, latency, and cost trade-offs that determine whether an ML product is viable Excellent problem-solving, debugging, and communication skills, including with non-technical stakeholders Preferred Qualifications: On-device or edge inference experience (Core ML, TensorFlow Lite, ExecuTorch) and integration into mobile applications Active learning or human-in-the-loop labeling systems Computer vision on small, long-tail, or industrial inspection datasets rather than large public benchmarks Experience with camera and sensor integration, or working alongside hardware teams on capture quality Experience with robotics, IoT, or edge computing (ROS or similar platforms) Familiarity with automotive service, dealership operations, or DMS ecosystems Contributions to open-source projects Why Join ATI: Be part of a groundbreaking startup transforming automotive service technology Work with a team of industry veterans and top-tier robotics and software talent Own the ML platform for a product that already has paying customers - your architecture decisions set the scaling ceiling Our customers are our investors, so you'll develop and test in real service lane environments A genuine data advantage: proprietary inspection imagery from real shops that no public dataset can replicate Clear Total Addressable Market with strong pull from B2B partners Competitive salary and comprehensive benefits package Prime location in Woburn, MA with on-site parking Collaborative, low-ego, high-intensity work environment

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board