it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

3 jobs found

Email me jobs like this
Refine Search
Current Search
senior lead ai engineer genai platform services
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Annapurna Labs (U.S.) Inc. Cupertino, California
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon's custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK, developed by the Annapurna Labs team at AWS, is the backbone for accelerating deep learning and GenAI workloads on Amazon's Inferentia and Trainium ML accelerators. This comprehensive toolkit includes an ML compiler, runtime, and application framework that seamlessly integrates with popular ML frameworks like PyTorch and JAX enabling unparalleled ML inference and training performance. The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML accelerators. Working across the stack from PyTorch till the hardware-software boundary, our engineers build systematic infrastructure, innovate new methods and create high-performance kernels for ML functions, ensuring every compute unit is fine tuned for optimal performance for our customers' demanding workloads. We combine deep hardware knowledge with ML expertise to push the boundaries of what's possible in AI acceleration. As part of the broader Neuron organization, our team works across multiple technology layers - from frameworks and kernels and collaborate with compiler to runtime and collectives. We not only optimize current performance but also contribute to future architecture designs, working closely with customers to enable their models and ensure optimal performance. This role offers a unique opportunity to work at the intersection of machine learning, high-performance computing, and distributed architectures, where you'll help shape the future of AI acceleration technology You will architect and implement business critical features, and mentor a brilliant team of experienced engineers. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. It is a very unique learning culture. The team works closely with customers on their model enablement, providing direct support and optimization expertise to ensure their machine learning workloads achieve optimal performance on AWS ML accelerators. The team collaborates with open source ecosystems to provide seamless integration and bring peak performance at scale for customers and developers. This role is responsible for development, enablement and performance tuning of a wide variety of LLM model families, including massive scale large language models like the Llama family, DeepSeek and beyond. The Inference Enablement and Acceleration team works side by side with compiler engineers and runtime engineers to create, build and tune distributed inference solutions with Trainium and Inferentia. Experience optimizing inference performance for both latency and throughput on such large models across the stack from system level optimizations through to Pytorch or JAX is a must have. You can learn more about Neuron Key job responsibilities This role will help lead the efforts in building distributed inference support for Pytorch in the Neuron SDK. This role will tune these models to ensure highest performance and maximize the efficiency of them running on the customer AWS Trainium and Inferentia silicon and servers. Strong software development using Python, System level programming and ML knowledge are both critical to this role. Our engineers collaborate across compiler, runtime, framework, and hardware teams to optimize machine learning workloads for our global customer base. Working at the intersection of software, hardware, and machine learning systems, you'll bring expertise in low-level optimization, system architecture, and ML model acceleration. In this role, you will: Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators. Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment. Build infrastructure to systematically analyze and onboard multiple models with diverse architecture. Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models Analyze and optimize system-level performance across multiple generations of Neuron hardware Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks Implement optimizations such as fusion, sharding, tiling, and scheduling Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines. Work directly with customers to enable and optimize their ML models on AWS accelerators Collaborate across teams to develop innovative optimization techniques A day in the life You will collaborate with a cross-functional team of applied scientists, system engineers, and product managers to deliver state-of-the-art inference capabilities for Generative AI applications. Your work will involve debugging performance issues, optimizing memory usage, and shaping the future of Neuron's inference stack across Amazon and the Open Source Community. As you design and code solutions to help our team drive efficiencies in software architecture, you'll create metrics, implement automation and other improvements, and resolve the root cause of software defects. You will also build high-impact solutions to deliver to our large customer base and participate in design discussions, code review, and communicate with internal and external stakeholders. You will work cross-functionally to help drive business decisions with your technical input. You will work in a startup-like development environment, where you're always working on the most important initiative. About the team The Inference Enablement and Acceleration team fosters a builder's culture where experimentation is encouraged, and impact is measurable. We emphasize collaboration, technical ownership, and continuous learning. Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future. Join us to solve some of the most interesting and impactful infrastructure challenges in AI/ML today. BASIC QUALIFICATIONS - Bachelor's degree in computer science or equivalent - 3+ years of non-internship professional software development experience - 3+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model execution. - Software development experience in C++, Python (experience in at least one language is required). - Strong understanding of system performance, memory management, and parallel computing principles. - Proficiency in debugging, profiling, and implementing best software engineering practices in large-scale systems. PREFERRED QUALIFICATIONS - Familiarity with PyTorch, JIT compilation, and AOT tracing. - Familiarity with CUDA kernels or equivalent ML or low-level kernels - Candidates with performant kernel development such as CUTLASS, FlashInfer etc., would be well suited. - Familiar with syntax and tile-level semantics similar to Triton. - Experience with online/offline inference serving with vLLM, SGLang, TensorRT or similar platforms in production environments. - Deep understanding of computer architecture, operation systems level software and working knowledge of parallel computing. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers . click apply for full job details
09/30/2026
Full time
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon's custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK, developed by the Annapurna Labs team at AWS, is the backbone for accelerating deep learning and GenAI workloads on Amazon's Inferentia and Trainium ML accelerators. This comprehensive toolkit includes an ML compiler, runtime, and application framework that seamlessly integrates with popular ML frameworks like PyTorch and JAX enabling unparalleled ML inference and training performance. The Inference Enablement and Acceleration team is at the forefront of running a wide range of models and supporting novel architecture alongside maximizing their performance for AWS's custom ML accelerators. Working across the stack from PyTorch till the hardware-software boundary, our engineers build systematic infrastructure, innovate new methods and create high-performance kernels for ML functions, ensuring every compute unit is fine tuned for optimal performance for our customers' demanding workloads. We combine deep hardware knowledge with ML expertise to push the boundaries of what's possible in AI acceleration. As part of the broader Neuron organization, our team works across multiple technology layers - from frameworks and kernels and collaborate with compiler to runtime and collectives. We not only optimize current performance but also contribute to future architecture designs, working closely with customers to enable their models and ensure optimal performance. This role offers a unique opportunity to work at the intersection of machine learning, high-performance computing, and distributed architectures, where you'll help shape the future of AI acceleration technology You will architect and implement business critical features, and mentor a brilliant team of experienced engineers. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. It is a very unique learning culture. The team works closely with customers on their model enablement, providing direct support and optimization expertise to ensure their machine learning workloads achieve optimal performance on AWS ML accelerators. The team collaborates with open source ecosystems to provide seamless integration and bring peak performance at scale for customers and developers. This role is responsible for development, enablement and performance tuning of a wide variety of LLM model families, including massive scale large language models like the Llama family, DeepSeek and beyond. The Inference Enablement and Acceleration team works side by side with compiler engineers and runtime engineers to create, build and tune distributed inference solutions with Trainium and Inferentia. Experience optimizing inference performance for both latency and throughput on such large models across the stack from system level optimizations through to Pytorch or JAX is a must have. You can learn more about Neuron Key job responsibilities This role will help lead the efforts in building distributed inference support for Pytorch in the Neuron SDK. This role will tune these models to ensure highest performance and maximize the efficiency of them running on the customer AWS Trainium and Inferentia silicon and servers. Strong software development using Python, System level programming and ML knowledge are both critical to this role. Our engineers collaborate across compiler, runtime, framework, and hardware teams to optimize machine learning workloads for our global customer base. Working at the intersection of software, hardware, and machine learning systems, you'll bring expertise in low-level optimization, system architecture, and ML model acceleration. In this role, you will: Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators. Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment. Build infrastructure to systematically analyze and onboard multiple models with diverse architecture. Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models Analyze and optimize system-level performance across multiple generations of Neuron hardware Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks Implement optimizations such as fusion, sharding, tiling, and scheduling Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines. Work directly with customers to enable and optimize their ML models on AWS accelerators Collaborate across teams to develop innovative optimization techniques A day in the life You will collaborate with a cross-functional team of applied scientists, system engineers, and product managers to deliver state-of-the-art inference capabilities for Generative AI applications. Your work will involve debugging performance issues, optimizing memory usage, and shaping the future of Neuron's inference stack across Amazon and the Open Source Community. As you design and code solutions to help our team drive efficiencies in software architecture, you'll create metrics, implement automation and other improvements, and resolve the root cause of software defects. You will also build high-impact solutions to deliver to our large customer base and participate in design discussions, code review, and communicate with internal and external stakeholders. You will work cross-functionally to help drive business decisions with your technical input. You will work in a startup-like development environment, where you're always working on the most important initiative. About the team The Inference Enablement and Acceleration team fosters a builder's culture where experimentation is encouraged, and impact is measurable. We emphasize collaboration, technical ownership, and continuous learning. Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future. Join us to solve some of the most interesting and impactful infrastructure challenges in AI/ML today. BASIC QUALIFICATIONS - Bachelor's degree in computer science or equivalent - 3+ years of non-internship professional software development experience - 3+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model execution. - Software development experience in C++, Python (experience in at least one language is required). - Strong understanding of system performance, memory management, and parallel computing principles. - Proficiency in debugging, profiling, and implementing best software engineering practices in large-scale systems. PREFERRED QUALIFICATIONS - Familiarity with PyTorch, JIT compilation, and AOT tracing. - Familiarity with CUDA kernels or equivalent ML or low-level kernels - Candidates with performant kernel development such as CUTLASS, FlashInfer etc., would be well suited. - Familiar with syntax and tile-level semantics similar to Triton. - Experience with online/offline inference serving with vLLM, SGLang, TensorRT or similar platforms in production environments. - Deep understanding of computer architecture, operation systems level software and working knowledge of parallel computing. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers . click apply for full job details
Principal AI Engineer
h2o.ai Addison, Texas
Job Description Job Description H2O.ai is on a mission to democratize AI for Good. As the world's leading agentic AI company, H2O.ai converges Generative and Predictive AI to help enterprises and public sector agencies develop purpose-built Agents, SLMs, and solutions on their private data. With a focus on secure, compliant, and infrastructure-flexible Sovereign AI deployments, H2O.ai delivers solutions that align with the highest standards of data privacy and control. Its open-source technology is trusted by over 20,000 organizations worldwide, including more than half of the Fortune 500. H2O.ai powers AI transformation for companies like AT&T, Commonwealth Bank of Australia, Wells Fargo, Bank of America, Workday, Progressive Insurance, and NIH. For more information, visit . About This Opportunity We are looking for a Principal AI Engineer who builds things that matter. You will design and ship end-to-end AI solutions for some of APAC's most complex enterprise problems - spanning agentic AI systems, LLM applications, and production ML pipelines. This is a hands-on engineering role embedded within a customer-facing field team, meaning your work will be seen, used, and evaluated by real enterprises from day one. You will work alongside Kaggle Grandmasters, ML engineers, and domain experts to deliver AI that goes beyond demos - into production, into workflows, and into measurable business outcomes. This position is based in Dallas, Texas and requires onsite customer interfacing. What You Will Do Customer Engagement Leadership Lead end-to-end technical engagement with enterprise customers, acting as the senior point of accountability for delivery quality, stakeholder relationships, and outcomes. Manage multiple concurrent engagement streams simultaneously - coordinating workplans, resourcing, and milestones across cross-functional teams. Serve as the primary technical escalation point for customer issues, proactively identifying risks and driving resolution across engineering, product, and leadership. Build and maintain trusted relationships with customer data science teams, engineering leads, and executive stakeholders - translating business needs into technical direction and back again. Lead pre-sales and proof-of-concept engagements, setting the technical strategy and ensuring the team delivers demonstrations that build genuine enterprise trust. Represent H2O.ai externally at customer workshops, executive briefings, and technical deep-dives as a credible senior voice. Agentic AI & LLM Engineering Design and build agentic AI systems and multi-agent frameworks that automate complex, multi-step enterprise workflows. Develop and deploy LLM-powered applications using RAG, fine-tuning, prompt engineering, function calling, and tool use. Implement guardrails, evaluation frameworks, and responsible AI controls to ensure production-grade reliability and safety. Stay current with the rapidly evolving agentic AI landscape - MCP, LLM orchestration frameworks, reasoning models - and bring the best into customer engagements. End-to-End AI Application Development Own the full development lifecycle across multiple streams: from problem framing and data exploration through model development, API integration, and production deployment. Build scalable backend services and APIs that expose AI capabilities to enterprise applications and workflows. Integrate AI models into customer environments - cloud, on-prem, and hybrid - ensuring performance, stability, and maintainability at scale. Develop ML pipelines and LLMOps infrastructure that support continuous model improvement and monitoring in production. Team Collaboration & Delivery Excellence Coordinate delivery across engineers, program managers, and solution architects - ensuring workstreams are aligned, unblocked, and progressing to plan. Set the technical bar for the engagements you lead, reviewing outputs, shaping architecture decisions, and ensuring engineering quality across the team. Mentor and guide junior ML engineers and solution engineers within engagements, building team capability alongside delivery. Collaborate closely with H2O.ai product and engineering teams to surface customer feedback, shape roadmap input, and resolve platform-level issues. What We Are Looking For Experience & Background 8+ years of hands-on AI/ML engineering experience, including end-to-end model development and production deployment. Demonstrable experience leading technical delivery across complex, multi-stakeholder enterprise engagements - not just executing within them. Demonstrable experience building LLM-powered applications - RAG pipelines, agentic workflows, fine-tuned models, or similar. Strong Python engineering skills; experience with ML frameworks (PyTorch, TensorFlow, scikit-learn) and LLM tooling (LangChain, LlamaIndex, or equivalent). Experience deploying AI services in cloud or enterprise environments (AWS, Azure, GCP, on-prem Kubernetes). Skills & Capabilities Proven ability to manage multiple concurrent workstreams and coordinate cross-functional teams toward shared delivery milestones. Deep understanding of modern GenAI concepts: prompt engineering, RAG, fine-tuning, RLHF, model evaluation, guardrails, and LLMOps. Solid grounding in classical ML - able to select the right tool for the problem, not just default to the latest LLM. Backend development skills: REST APIs, containerisation (Docker/Kubernetes), and CI/CD pipelines for AI applications. Strong executive communication - able to run a board-level briefing one hour and a technical design review the next, credibly. Comfortable with ambiguity and able to set direction for a team when requirements are incomplete or evolving. How to Stand Out From the Crowd Kaggle or competitive ML experience. Familiarity with H2O.ai products, Wave, or H2O Document AI. Experience in financial services, healthcare, or other regulated industry AI deployments. Exposure to tabular foundation models, AutoML, or enterprise ML platforms. Prior experience in a customer-facing or field engineering role. Why H2O.ai? Market leader in total rewards Remote-friendly culture Flexible working environment Be part of a world-class team Career growth The base salary for this role ranges from $175,000 to $200,000. Compensation is determined based on several factors, including skills, experience, job scope, location, and relevant market compensation data. H2O.ai is committed to creating a diverse and inclusive culture. All qualified applicants will receive consideration for employment without regard to their race, ethnicity, religion, gender, sexual orientation, age, disability status or any other legally protected basis. H2O.ai is an innovative AI cloud platform company, leading the mission to democratize AI for everyone. Thousands of organizations from all over the world have used our cutting-edge technology across a variety of industries. We've made it easy for people at all levels to generate breakthrough solutions to complex business problems and advance the discovery of new ideas and revenue streams. We push the boundaries of what is possible with artificial intelligence. H2O.ai employs the world's top Kaggle Grandmasters, the community of best-in-the-world machine learning practitioners and data scientists. A strong AI for Good ethos and responsible AI drive the company's purpose. Please visit to learn more.
09/28/2026
Full time
Job Description Job Description H2O.ai is on a mission to democratize AI for Good. As the world's leading agentic AI company, H2O.ai converges Generative and Predictive AI to help enterprises and public sector agencies develop purpose-built Agents, SLMs, and solutions on their private data. With a focus on secure, compliant, and infrastructure-flexible Sovereign AI deployments, H2O.ai delivers solutions that align with the highest standards of data privacy and control. Its open-source technology is trusted by over 20,000 organizations worldwide, including more than half of the Fortune 500. H2O.ai powers AI transformation for companies like AT&T, Commonwealth Bank of Australia, Wells Fargo, Bank of America, Workday, Progressive Insurance, and NIH. For more information, visit . About This Opportunity We are looking for a Principal AI Engineer who builds things that matter. You will design and ship end-to-end AI solutions for some of APAC's most complex enterprise problems - spanning agentic AI systems, LLM applications, and production ML pipelines. This is a hands-on engineering role embedded within a customer-facing field team, meaning your work will be seen, used, and evaluated by real enterprises from day one. You will work alongside Kaggle Grandmasters, ML engineers, and domain experts to deliver AI that goes beyond demos - into production, into workflows, and into measurable business outcomes. This position is based in Dallas, Texas and requires onsite customer interfacing. What You Will Do Customer Engagement Leadership Lead end-to-end technical engagement with enterprise customers, acting as the senior point of accountability for delivery quality, stakeholder relationships, and outcomes. Manage multiple concurrent engagement streams simultaneously - coordinating workplans, resourcing, and milestones across cross-functional teams. Serve as the primary technical escalation point for customer issues, proactively identifying risks and driving resolution across engineering, product, and leadership. Build and maintain trusted relationships with customer data science teams, engineering leads, and executive stakeholders - translating business needs into technical direction and back again. Lead pre-sales and proof-of-concept engagements, setting the technical strategy and ensuring the team delivers demonstrations that build genuine enterprise trust. Represent H2O.ai externally at customer workshops, executive briefings, and technical deep-dives as a credible senior voice. Agentic AI & LLM Engineering Design and build agentic AI systems and multi-agent frameworks that automate complex, multi-step enterprise workflows. Develop and deploy LLM-powered applications using RAG, fine-tuning, prompt engineering, function calling, and tool use. Implement guardrails, evaluation frameworks, and responsible AI controls to ensure production-grade reliability and safety. Stay current with the rapidly evolving agentic AI landscape - MCP, LLM orchestration frameworks, reasoning models - and bring the best into customer engagements. End-to-End AI Application Development Own the full development lifecycle across multiple streams: from problem framing and data exploration through model development, API integration, and production deployment. Build scalable backend services and APIs that expose AI capabilities to enterprise applications and workflows. Integrate AI models into customer environments - cloud, on-prem, and hybrid - ensuring performance, stability, and maintainability at scale. Develop ML pipelines and LLMOps infrastructure that support continuous model improvement and monitoring in production. Team Collaboration & Delivery Excellence Coordinate delivery across engineers, program managers, and solution architects - ensuring workstreams are aligned, unblocked, and progressing to plan. Set the technical bar for the engagements you lead, reviewing outputs, shaping architecture decisions, and ensuring engineering quality across the team. Mentor and guide junior ML engineers and solution engineers within engagements, building team capability alongside delivery. Collaborate closely with H2O.ai product and engineering teams to surface customer feedback, shape roadmap input, and resolve platform-level issues. What We Are Looking For Experience & Background 8+ years of hands-on AI/ML engineering experience, including end-to-end model development and production deployment. Demonstrable experience leading technical delivery across complex, multi-stakeholder enterprise engagements - not just executing within them. Demonstrable experience building LLM-powered applications - RAG pipelines, agentic workflows, fine-tuned models, or similar. Strong Python engineering skills; experience with ML frameworks (PyTorch, TensorFlow, scikit-learn) and LLM tooling (LangChain, LlamaIndex, or equivalent). Experience deploying AI services in cloud or enterprise environments (AWS, Azure, GCP, on-prem Kubernetes). Skills & Capabilities Proven ability to manage multiple concurrent workstreams and coordinate cross-functional teams toward shared delivery milestones. Deep understanding of modern GenAI concepts: prompt engineering, RAG, fine-tuning, RLHF, model evaluation, guardrails, and LLMOps. Solid grounding in classical ML - able to select the right tool for the problem, not just default to the latest LLM. Backend development skills: REST APIs, containerisation (Docker/Kubernetes), and CI/CD pipelines for AI applications. Strong executive communication - able to run a board-level briefing one hour and a technical design review the next, credibly. Comfortable with ambiguity and able to set direction for a team when requirements are incomplete or evolving. How to Stand Out From the Crowd Kaggle or competitive ML experience. Familiarity with H2O.ai products, Wave, or H2O Document AI. Experience in financial services, healthcare, or other regulated industry AI deployments. Exposure to tabular foundation models, AutoML, or enterprise ML platforms. Prior experience in a customer-facing or field engineering role. Why H2O.ai? Market leader in total rewards Remote-friendly culture Flexible working environment Be part of a world-class team Career growth The base salary for this role ranges from $175,000 to $200,000. Compensation is determined based on several factors, including skills, experience, job scope, location, and relevant market compensation data. H2O.ai is committed to creating a diverse and inclusive culture. All qualified applicants will receive consideration for employment without regard to their race, ethnicity, religion, gender, sexual orientation, age, disability status or any other legally protected basis. H2O.ai is an innovative AI cloud platform company, leading the mission to democratize AI for everyone. Thousands of organizations from all over the world have used our cutting-edge technology across a variety of industries. We've made it easy for people at all levels to generate breakthrough solutions to complex business problems and advance the discovery of new ideas and revenue streams. We push the boundaries of what is possible with artificial intelligence. H2O.ai employs the world's top Kaggle Grandmasters, the community of best-in-the-world machine learning practitioners and data scientists. A strong AI for Good ethos and responsible AI drive the company's purpose. Please visit to learn more.
KPMG
Lead Specialist, AI Solution Architect
KPMG Los Angeles, California
The KPMG Advisory practice is at the forefront of transformation, offering excellent opportunities for individuals to advance their careers and expertise with KPMG. Looking ahead, we anticipate continued evolution and success within the practice, fostering both personal and professional development, thereby creating new pathways for growth. In this ever-changing market environment, our professionals must be adaptable and thrive in a collaborative, team-driven culture. At KPMG, our people are our number one priority. With a wealth of learning and career development opportunities, a world-class training facility, and leading market tools, we help our people continue to grow both professionally and personally. If you're looking for a firm with a strong team connection where you can be your whole self, have an impact, advance your skills, deepen your experiences, and have the flexibility and access to constantly find new areas of inspiration and expand your capabilities, then consider a career in Advisory. KPMG is currently seeking a Lead Specialist, AI Solution Architect to join our KPMG Managed Services practice. Responsibilities: Architect and lead end-to-end delivery of enterprise-scale AI solutions using Agile and DevOps practices, providing technical leadership across planning, development, code quality, reviews, and release management with tools such as Azure DevOps, JIRA, Git, and CI/CD pipelines. Design and implement secure, resilient, and scalable cloud-native architectures on Microsoft Azure, leveraging IaaS/PaaS services, modern application stacks, and enterprise data platforms to meet regulatory and performance requirements. Drive AI, GenAI, and agent-based solution strategy by leading proofs of concept and pilots, and guiding successful initiatives through transition into production-ready, operational platforms. Translate complex business needs into pragmatic technical solutions by partnering with business stakeholders, product owners, and enterprise architects to shape roadmaps, evaluate emerging technologies, and align AI capabilities to measurable business outcomes. Design and enable advanced agentic and intelligent automation patterns, including retrieval-augmented generation (RAG), multi-agent and A2A architectures, orchestration, state management, observability, and interoperable workflows that scale beyond pilot stages. Establish and enforce standards for architecture artifacts, documentation, deployment patterns, monitoring, and support while embedding security, privacy, and responsible AI principles into all solution designs; mentor and guide onshore and offshore engineering teams to foster a high-performance, innovation-driven culture. Act with integrity, professionalism, and personal responsibility to uphold KPMG's respectful and courteous work environment. Qualifications: Minimum of 5 years of experience designing and leading enterprise-scale technology solutions, with demonstrated ownership of architecture and delivery in complex environments. Bachelor's degree from an accredited college or university in Computer Science, Engineering, or a related field is required. Strong background in cloud-native architecture on Microsoft Azure, including application services, data platforms, integration patterns, and DevOps automation. Proven expertise in AI-driven and modern data architectures, including GenAI, agent-based systems, and RAG-style solutions, with experience taking concepts through production deployment. Solid understanding of full-stack systems, including modern web frameworks, backend services, APIs, and data integrations, with the ability to make pragmatic architectural trade-offs. Demonstrated experience leading and mentoring technical teams, influencing senior stakeholders, and serving as a trusted technical advisor with strong written and verbal communication skills. Ability to travel as required. Applicants must be authorized to work in the U.S. without the need for employment based visa sponsorship now or in the future. KPMG LLP will not sponsor applicants for U.S. work visa status for this opportunity (no sponsorship is available for H-1B, L-1, TN, O-1, E-3, H-1B1, F-1, J-1, OPT, CPT or any other employment based visa). KPMG LLP and its affiliates and subsidiaries ("KPMG") complies with all local/state regulations regarding displaying salary ranges. If required, the ranges displayed below or via the URL below are specifically for those potential hires who will work in the location(s) listed. Any offered salary is determined based on relevant factors such as applicant's skills, job responsibilities, prior relevant experience, certain degrees and certifications and market considerations. In addition, KPMG is proud to offer a comprehensive, competitive benefits package, with options designed to help you make the best decisions for yourself, your family, and your lifestyle. Available benefits are based on eligibility. Our Total Rewards package includes a variety of medical and dental plans, vision coverage, disability and life insurance, 401(k) plans, and a robust suite of personal well-being benefits to support your mental health. Depending on job classification, standard work hours, and years of service, KPMG provides Personal Time Off per fiscal year. Additionally, each year KPMG publishes a calendar of holidays to be observed during the year and provides eligible employees two breaks each year where employees will not be required to use Personal Time Off; one is at year end and the other is around the July 4th holiday. Additional details about our benefits can be found towards the bottom of our KPMG US Careers site at Benefits & How We Work . Follow this link to obtain salary ranges by city outside of CA: California Salary Range: $124735 - $254495 KPMG offers a comprehensive compensation and benefits package. KPMG is an equal opportunity employer. KPMG complies with all applicable federal, state and local laws regarding recruitment and hiring. All qualified applicants are considered for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, citizenship status, disability, protected veteran status, or any other category protected by applicable federal, state, or local laws. The attached link contains further information regarding KPMG's compliance with federal, state and local recruitment and hiring laws. No phone calls or agencies please. KPMG recruits on a rolling basis. Candidates are considered as they apply, until the opportunity is filled. Candidates are encouraged to apply expeditiously to any role(s) for which they are qualified that is also of interest to them. Los Angeles County applicants: Material job duties for this position are listed above. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, and safeguard business operations and company reputation. Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
07/14/2026
Full time
The KPMG Advisory practice is at the forefront of transformation, offering excellent opportunities for individuals to advance their careers and expertise with KPMG. Looking ahead, we anticipate continued evolution and success within the practice, fostering both personal and professional development, thereby creating new pathways for growth. In this ever-changing market environment, our professionals must be adaptable and thrive in a collaborative, team-driven culture. At KPMG, our people are our number one priority. With a wealth of learning and career development opportunities, a world-class training facility, and leading market tools, we help our people continue to grow both professionally and personally. If you're looking for a firm with a strong team connection where you can be your whole self, have an impact, advance your skills, deepen your experiences, and have the flexibility and access to constantly find new areas of inspiration and expand your capabilities, then consider a career in Advisory. KPMG is currently seeking a Lead Specialist, AI Solution Architect to join our KPMG Managed Services practice. Responsibilities: Architect and lead end-to-end delivery of enterprise-scale AI solutions using Agile and DevOps practices, providing technical leadership across planning, development, code quality, reviews, and release management with tools such as Azure DevOps, JIRA, Git, and CI/CD pipelines. Design and implement secure, resilient, and scalable cloud-native architectures on Microsoft Azure, leveraging IaaS/PaaS services, modern application stacks, and enterprise data platforms to meet regulatory and performance requirements. Drive AI, GenAI, and agent-based solution strategy by leading proofs of concept and pilots, and guiding successful initiatives through transition into production-ready, operational platforms. Translate complex business needs into pragmatic technical solutions by partnering with business stakeholders, product owners, and enterprise architects to shape roadmaps, evaluate emerging technologies, and align AI capabilities to measurable business outcomes. Design and enable advanced agentic and intelligent automation patterns, including retrieval-augmented generation (RAG), multi-agent and A2A architectures, orchestration, state management, observability, and interoperable workflows that scale beyond pilot stages. Establish and enforce standards for architecture artifacts, documentation, deployment patterns, monitoring, and support while embedding security, privacy, and responsible AI principles into all solution designs; mentor and guide onshore and offshore engineering teams to foster a high-performance, innovation-driven culture. Act with integrity, professionalism, and personal responsibility to uphold KPMG's respectful and courteous work environment. Qualifications: Minimum of 5 years of experience designing and leading enterprise-scale technology solutions, with demonstrated ownership of architecture and delivery in complex environments. Bachelor's degree from an accredited college or university in Computer Science, Engineering, or a related field is required. Strong background in cloud-native architecture on Microsoft Azure, including application services, data platforms, integration patterns, and DevOps automation. Proven expertise in AI-driven and modern data architectures, including GenAI, agent-based systems, and RAG-style solutions, with experience taking concepts through production deployment. Solid understanding of full-stack systems, including modern web frameworks, backend services, APIs, and data integrations, with the ability to make pragmatic architectural trade-offs. Demonstrated experience leading and mentoring technical teams, influencing senior stakeholders, and serving as a trusted technical advisor with strong written and verbal communication skills. Ability to travel as required. Applicants must be authorized to work in the U.S. without the need for employment based visa sponsorship now or in the future. KPMG LLP will not sponsor applicants for U.S. work visa status for this opportunity (no sponsorship is available for H-1B, L-1, TN, O-1, E-3, H-1B1, F-1, J-1, OPT, CPT or any other employment based visa). KPMG LLP and its affiliates and subsidiaries ("KPMG") complies with all local/state regulations regarding displaying salary ranges. If required, the ranges displayed below or via the URL below are specifically for those potential hires who will work in the location(s) listed. Any offered salary is determined based on relevant factors such as applicant's skills, job responsibilities, prior relevant experience, certain degrees and certifications and market considerations. In addition, KPMG is proud to offer a comprehensive, competitive benefits package, with options designed to help you make the best decisions for yourself, your family, and your lifestyle. Available benefits are based on eligibility. Our Total Rewards package includes a variety of medical and dental plans, vision coverage, disability and life insurance, 401(k) plans, and a robust suite of personal well-being benefits to support your mental health. Depending on job classification, standard work hours, and years of service, KPMG provides Personal Time Off per fiscal year. Additionally, each year KPMG publishes a calendar of holidays to be observed during the year and provides eligible employees two breaks each year where employees will not be required to use Personal Time Off; one is at year end and the other is around the July 4th holiday. Additional details about our benefits can be found towards the bottom of our KPMG US Careers site at Benefits & How We Work . Follow this link to obtain salary ranges by city outside of CA: California Salary Range: $124735 - $254495 KPMG offers a comprehensive compensation and benefits package. KPMG is an equal opportunity employer. KPMG complies with all applicable federal, state and local laws regarding recruitment and hiring. All qualified applicants are considered for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, citizenship status, disability, protected veteran status, or any other category protected by applicable federal, state, or local laws. The attached link contains further information regarding KPMG's compliance with federal, state and local recruitment and hiring laws. No phone calls or agencies please. KPMG recruits on a rolling basis. Candidates are considered as they apply, until the opportunity is filled. Candidates are encouraged to apply expeditiously to any role(s) for which they are qualified that is also of interest to them. Los Angeles County applicants: Material job duties for this position are listed above. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, and safeguard business operations and company reputation. Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board