it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice

Modal title

11 jobs found in Cupertino

Software Development Manager - NKI Compiler, AWS Neuron, Annapurna Labs
Annapurna Labs (U.S.) Inc. Cupertino, California
The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. Trainium delivers best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, the Neuron Kernel Interface (NKI) compiler, and a runtime that natively integrates into popular ML frameworks such as PyTorch and JAX. Neuron Kernel Interface (NKI) is a bare-metal language and compiler for directly programming NeuronDevices available on AWS Trainium instances. You can use NKI to develop, optimize, and run new operators directly on NeuronCores while making full use of available compute and memory resources. Explore NKI: - AWS Neuron is used at scale by customers such as Epic Games, Snap, Airbnb, Autodesk, Amazon Alexa, and Amazon Rekognition, along with many others across a range of segments. The Team: The Amazon Annapurna Labs team is responsible for building innovative silicon and software for AWS customers. We are at the forefront of innovation, combining cloud scale with the world's most talented engineers. Our team covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. With such breadth of talent, there is opportunity to learn all of the time. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. When you couple that with the ability to work on so many different products and services, it makes for a unique learning culture. Learn more about our history: You: We are seeking a talented Software Development Manager with strong leadership and mentoring skills to join our NKI development team. As an SDM III, you will lead a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You will need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron manager, innovating on behalf of our customers. You will draw on knowledge of resource management, scheduling, code generation, optimization, and instruction architectures across CPU, NPU, GPU, and novel forms of compute. You will leverage your technical communication skills to partner with AWS ML services teams and pre-silicon design, and to bring new products and features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that boosts the entire deep learning community. Explore the Product: - - In order to be considered for this role, candidates must be currently located in or willing to relocate to Cupertino, CA. Key job responsibilities - Lead, grow, and mentor a team of compiler engineers, including hiring, career development, and team health - Own execution and delivery of NKI compiler features across release cycles, balancing scope, quality, and timelines - Set technical direction and roadmap for your area of the NKI compiler in partnership with senior engineers - Partner across teams, including frameworks, kernel development, runtime, hardware/pre-silicon design, and product management, to bring new features to market - Drive resolution of complex, high-priority technical issues, staying close enough to the code and hardware to guide the team - Represent your team's work and priorities to senior leadership and cross-organizational stakeholders - Anchor decisions in customer impact, ensuring the team builds what unblocks real ML workloads on Trainium A day in the life No two days look the same on the NKI team, but most blend people leadership, technical depth, and cross-team collaboration. You might start the morning in a 1:1 with an engineer, working through the design of a new compiler optimization or unblocking a tricky scheduling problem. Mid-morning, you join a release sync to check in on branch-cut readiness and make a call on what makes the current train versus the next. After lunch, you catch up with a partner team, kernel developers, runtime, or hardware design, to align on an upcoming feature and its dependencies. In the afternoon, you spend focused time on the roadmap: shaping where NKI is heading over the next few quarters, and translating customer needs into concrete technical investments. You review a design doc, leave feedback that sharpens the team's thinking, and dig into a profiling result yourself to understand where a real workload is leaving performance on the table. You close the day by clearing a path for your team and resolving an escalation, connecting two people who should be talking, or writing up a decision so the team can move fast tomorrow. Throughout, you keep one question at the center: what does this unblock for our customers? You are technically credible enough to earn your team's trust, and you spend your energy on the highest-leverage problems: growing your people, delivering the compiler, and inventing on behalf of the customers who run their most demanding ML workloads on Trainium. About the team Inclusive Team Culture Here at Annapurna Labs, we embrace our differences. We are committed to furthering our culture of inclusion. Amazon has ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS- 2+ years of engineering team management experience - 6+ years of working directly within engineering teams experience - 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - Understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization) - Strong software design fundamentals and excellent system-level coding skills PREFERRED QUALIFICATIONS- M.S. or Ph.D. in Computer Science or related technical field Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line . click apply for full job details
08/17/2026
Full time
The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. Trainium delivers best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, the Neuron Kernel Interface (NKI) compiler, and a runtime that natively integrates into popular ML frameworks such as PyTorch and JAX. Neuron Kernel Interface (NKI) is a bare-metal language and compiler for directly programming NeuronDevices available on AWS Trainium instances. You can use NKI to develop, optimize, and run new operators directly on NeuronCores while making full use of available compute and memory resources. Explore NKI: - AWS Neuron is used at scale by customers such as Epic Games, Snap, Airbnb, Autodesk, Amazon Alexa, and Amazon Rekognition, along with many others across a range of segments. The Team: The Amazon Annapurna Labs team is responsible for building innovative silicon and software for AWS customers. We are at the forefront of innovation, combining cloud scale with the world's most talented engineers. Our team covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. With such breadth of talent, there is opportunity to learn all of the time. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. When you couple that with the ability to work on so many different products and services, it makes for a unique learning culture. Learn more about our history: You: We are seeking a talented Software Development Manager with strong leadership and mentoring skills to join our NKI development team. As an SDM III, you will lead a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You will need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron manager, innovating on behalf of our customers. You will draw on knowledge of resource management, scheduling, code generation, optimization, and instruction architectures across CPU, NPU, GPU, and novel forms of compute. You will leverage your technical communication skills to partner with AWS ML services teams and pre-silicon design, and to bring new products and features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that boosts the entire deep learning community. Explore the Product: - - In order to be considered for this role, candidates must be currently located in or willing to relocate to Cupertino, CA. Key job responsibilities - Lead, grow, and mentor a team of compiler engineers, including hiring, career development, and team health - Own execution and delivery of NKI compiler features across release cycles, balancing scope, quality, and timelines - Set technical direction and roadmap for your area of the NKI compiler in partnership with senior engineers - Partner across teams, including frameworks, kernel development, runtime, hardware/pre-silicon design, and product management, to bring new features to market - Drive resolution of complex, high-priority technical issues, staying close enough to the code and hardware to guide the team - Represent your team's work and priorities to senior leadership and cross-organizational stakeholders - Anchor decisions in customer impact, ensuring the team builds what unblocks real ML workloads on Trainium A day in the life No two days look the same on the NKI team, but most blend people leadership, technical depth, and cross-team collaboration. You might start the morning in a 1:1 with an engineer, working through the design of a new compiler optimization or unblocking a tricky scheduling problem. Mid-morning, you join a release sync to check in on branch-cut readiness and make a call on what makes the current train versus the next. After lunch, you catch up with a partner team, kernel developers, runtime, or hardware design, to align on an upcoming feature and its dependencies. In the afternoon, you spend focused time on the roadmap: shaping where NKI is heading over the next few quarters, and translating customer needs into concrete technical investments. You review a design doc, leave feedback that sharpens the team's thinking, and dig into a profiling result yourself to understand where a real workload is leaving performance on the table. You close the day by clearing a path for your team and resolving an escalation, connecting two people who should be talking, or writing up a decision so the team can move fast tomorrow. Throughout, you keep one question at the center: what does this unblock for our customers? You are technically credible enough to earn your team's trust, and you spend your energy on the highest-leverage problems: growing your people, delivering the compiler, and inventing on behalf of the customers who run their most demanding ML workloads on Trainium. About the team Inclusive Team Culture Here at Annapurna Labs, we embrace our differences. We are committed to furthering our culture of inclusion. Amazon has ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS- 2+ years of engineering team management experience - 6+ years of working directly within engineering teams experience - 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - Understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization) - Strong software design fundamentals and excellent system-level coding skills PREFERRED QUALIFICATIONS- M.S. or Ph.D. in Computer Science or related technical field Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line . click apply for full job details
Software Development Manager, AWS Neuron SDK - Distributed Training
Amazon Development Center U.S., Inc. Cupertino, California
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trainium-based servers that use them. As the SDM of Software Development for the Neuron Training team, you will be responsible for leading a strong team of engineers and managers to help design and deploy these new products. A successful candidate will have an established background in developing Machine Learning products with direct customer-facing experience, a strong technical ability and a motivation to achieve results. Experience in Machine Learning and software development is also a must. Responsible for the full development life cycle of our integrations and extensions for training support in Pytorch, JAX, and distributed training libraries with a focus on performance of latest ML models at scale on Trainium using latest techniques in performance optimization, accuracy, and resilience. you will lead the way to ensure support for key ML functionality in a combined chip / software platform, which will ensure the right thing is being built and delivered to customers Key job responsibilities Lead a team of engineers focused on enabling new ML training customers on the Neuron SDK / Trainium platform. Own the customer onboarding journey from model evaluation through production training at scale Drive engineering initiatives to maximize Model FLOPS Utilization (MFU) for customer workloads through performance analysis, profiling, and tuning tools. Build and maintain tooling, automation, and documentation that accelerates time-to-first-training for new customer models. Partner with compiler, runtime, and framework teams to identify and resolve blockers in customer workloads. Develop scalable processes for distributed training enablement (FSDP, DeepSpeed, Megatron, custom parallelism strategies). Drive technical strategy for supporting frontier model architectures (LLMs, MoE, multi-modal) on Trainium. Build strong cross-functional partnerships with product management, developer relations, and customer-facing teams. Recruit, mentor, and grow a high-performing team of ML systems engineers. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team About Us Inclusive Team Culture Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS- Experience working with PyTorch or JAX software - 3+ years of engineering team management experience - 7+ years of working directly within engineering teams experience - 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - 3+ years of experience in deep learning / machine learning, including model training workflows - Experience with distributed training at scale (multi-node, multi-accelerator) PREFERRED QUALIFICATIONS- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience directly managing scientists or machine learning engineers - Experience debugging, profiling, and implementing best software engineering practices in large-scale systems - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels - Experience with performance analysis, profiling, and optimization for deep learning training workloads Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 212 700.00 USD annually
08/17/2026
Full time
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trainium-based servers that use them. As the SDM of Software Development for the Neuron Training team, you will be responsible for leading a strong team of engineers and managers to help design and deploy these new products. A successful candidate will have an established background in developing Machine Learning products with direct customer-facing experience, a strong technical ability and a motivation to achieve results. Experience in Machine Learning and software development is also a must. Responsible for the full development life cycle of our integrations and extensions for training support in Pytorch, JAX, and distributed training libraries with a focus on performance of latest ML models at scale on Trainium using latest techniques in performance optimization, accuracy, and resilience. you will lead the way to ensure support for key ML functionality in a combined chip / software platform, which will ensure the right thing is being built and delivered to customers Key job responsibilities Lead a team of engineers focused on enabling new ML training customers on the Neuron SDK / Trainium platform. Own the customer onboarding journey from model evaluation through production training at scale Drive engineering initiatives to maximize Model FLOPS Utilization (MFU) for customer workloads through performance analysis, profiling, and tuning tools. Build and maintain tooling, automation, and documentation that accelerates time-to-first-training for new customer models. Partner with compiler, runtime, and framework teams to identify and resolve blockers in customer workloads. Develop scalable processes for distributed training enablement (FSDP, DeepSpeed, Megatron, custom parallelism strategies). Drive technical strategy for supporting frontier model architectures (LLMs, MoE, multi-modal) on Trainium. Build strong cross-functional partnerships with product management, developer relations, and customer-facing teams. Recruit, mentor, and grow a high-performing team of ML systems engineers. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team About Us Inclusive Team Culture Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS- Experience working with PyTorch or JAX software - 3+ years of engineering team management experience - 7+ years of working directly within engineering teams experience - 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - 3+ years of experience in deep learning / machine learning, including model training workflows - Experience with distributed training at scale (multi-node, multi-accelerator) PREFERRED QUALIFICATIONS- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience directly managing scientists or machine learning engineers - Experience debugging, profiling, and implementing best software engineering practices in large-scale systems - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels - Experience with performance analysis, profiling, and optimization for deep learning training workloads Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 212 700.00 USD annually
Software Engineer Neuron Containers, Annapurna Labs, Neuron Containers, Annapurna Labs
Annapurna Labs (U.S.) Inc. Cupertino, California
Annapurna Labs was a startup acquired by AWS in 2015 and is now fully integrated. AWS Neuron is the complete software stack for the Inferentia and Trainium ML accelerators - the custom silicon that powers large-scale AI workloads on AWS. The Neuron Containers team is looking for a Software Development Engineer to build platform integrations that enable customers to run distributed training and inference workloads on Neuron at scale. The team owns Neuron integration with Kubernetes, ECS, and Slurm - handling device allocation, fault tolerance, auto-scaling, and orchestration across large clusters. The team also owns delivery of Neuron Deep Learning Containers (DLCs) and Deep Learning AMIs (DLAMIs) - the pre-configured container images and machine images that package the Neuron SDK for customer deployment on Trainium and Inferentia instances. Key job responsibilities Design and implement container platform integrations - device plugins, DRA drivers, and operators for ML accelerator resource management Build and maintain Neuron DLCs and DLAMIs for customer deployment across EKS, ECS, EC2, and SageMaker Diagnose and resolve performance and scalability issues across large customer clusters Simplify systems - deprecate legacy software and reduce complexity in container delivery pipelines Deliver software across the full development lifecycle including design documentation, implementation, testing, deployment, and operations A day in the life You'll work with teams across the Neuron org and customers to build and maintain integrations for current and next-generation accelerators. You'll participate in architecture reviews, triage test failures, resolve operational issues, and contribute to upstream Kubernetes projects. You'll debug platform integration problems - how Neuron interacts with container runtimes, orchestrators, and scheduling systems a scale. BASIC QUALIFICATIONS- 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent - Experience with distributed systems or large-scale cluster infrastructure - Familiarity with ML training/inference workflows (distributed training, collective - Experience with AWS compute services (EC2, EKS, ECS, ECR) - Hands-on experience with Helm, Prometheus, or Kubernetes operator frameworks - Experience with container image pipelines, Deep Learning Containers, or Deep Learning AMIs - Contributions to open-source projects, particularly in the Kubernetes ecosystem Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually
08/17/2026
Full time
Annapurna Labs was a startup acquired by AWS in 2015 and is now fully integrated. AWS Neuron is the complete software stack for the Inferentia and Trainium ML accelerators - the custom silicon that powers large-scale AI workloads on AWS. The Neuron Containers team is looking for a Software Development Engineer to build platform integrations that enable customers to run distributed training and inference workloads on Neuron at scale. The team owns Neuron integration with Kubernetes, ECS, and Slurm - handling device allocation, fault tolerance, auto-scaling, and orchestration across large clusters. The team also owns delivery of Neuron Deep Learning Containers (DLCs) and Deep Learning AMIs (DLAMIs) - the pre-configured container images and machine images that package the Neuron SDK for customer deployment on Trainium and Inferentia instances. Key job responsibilities Design and implement container platform integrations - device plugins, DRA drivers, and operators for ML accelerator resource management Build and maintain Neuron DLCs and DLAMIs for customer deployment across EKS, ECS, EC2, and SageMaker Diagnose and resolve performance and scalability issues across large customer clusters Simplify systems - deprecate legacy software and reduce complexity in container delivery pipelines Deliver software across the full development lifecycle including design documentation, implementation, testing, deployment, and operations A day in the life You'll work with teams across the Neuron org and customers to build and maintain integrations for current and next-generation accelerators. You'll participate in architecture reviews, triage test failures, resolve operational issues, and contribute to upstream Kubernetes projects. You'll debug platform integration problems - how Neuron interacts with container runtimes, orchestrators, and scheduling systems a scale. BASIC QUALIFICATIONS- 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent - Experience with distributed systems or large-scale cluster infrastructure - Familiarity with ML training/inference workflows (distributed training, collective - Experience with AWS compute services (EC2, EKS, ECS, ECR) - Hands-on experience with Helm, Prometheus, or Kubernetes operator frameworks - Experience with container image pipelines, Deep Learning Containers, or Deep Learning AMIs - Contributions to open-source projects, particularly in the Kubernetes ecosystem Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually
Neuron Runtime Software Development Engineer , Neuron Runtime
Annapurna Labs (U.S.) Inc. Cupertino, California
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the servers that use them. As the Software Development Engineer for the Neuron Runtime Team, you will be responsible for working alongside a team of engineers to develop and maintain high-performance runtime libraries and drivers for machine learning applications and AI accelerators. You will work on design, development, and deployment of Neuron Runtime and other Neuron components. The profiler plays a crucial role to internal and external customers in optimizing AI workloads across hardware platforms such as Trainium and Inferentia devices, by providing deep insights into performance bottlenecks and system behavior. Improving performance of ML Kernels and ML Frameworks. In this role, you will manage the full development life cycle of the Neuron Runtime, ensuring scalability, reliability, and usability. You will collaborate with cross-functional teams to ensure that the our C++ compiler generates key information so customers can understand and optimize the performance of our custom hardware. Additionally, you will drive innovations that allow the profiler to support multiple frameworks, such as PyTorch, JAX, and XLA. A successful candidate will have experience in architecting, building, and operating distributed systems with a focus on high availability and fault tolerance, Hands-on experience with AWS services (e.g., EC2, ECS, CloudWatch, S3, Lambda) in production environments and track record in Owning services end-to-end including deployment, monitoring, alarming, on-call, and post-incident review. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS- 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually USA, WA, Seattle - 143 400.00 USD annually
08/17/2026
Full time
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the servers that use them. As the Software Development Engineer for the Neuron Runtime Team, you will be responsible for working alongside a team of engineers to develop and maintain high-performance runtime libraries and drivers for machine learning applications and AI accelerators. You will work on design, development, and deployment of Neuron Runtime and other Neuron components. The profiler plays a crucial role to internal and external customers in optimizing AI workloads across hardware platforms such as Trainium and Inferentia devices, by providing deep insights into performance bottlenecks and system behavior. Improving performance of ML Kernels and ML Frameworks. In this role, you will manage the full development life cycle of the Neuron Runtime, ensuring scalability, reliability, and usability. You will collaborate with cross-functional teams to ensure that the our C++ compiler generates key information so customers can understand and optimize the performance of our custom hardware. Additionally, you will drive innovations that allow the profiler to support multiple frameworks, such as PyTorch, JAX, and XLA. A successful candidate will have experience in architecting, building, and operating distributed systems with a focus on high availability and fault tolerance, Hands-on experience with AWS services (e.g., EC2, ECS, CloudWatch, S3, Lambda) in production environments and track record in Owning services end-to-end including deployment, monitoring, alarming, on-call, and post-incident review. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS- 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually USA, WA, Seattle - 143 400.00 USD annually
Software Development Manager, AWS Neuron SDK - Distributed Training
Amazon Development Center U.S., Inc. Cupertino, California
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trainium-based servers that use them. As the SDM of Software Development for the Neuron Training team, you will be responsible for leading a strong team of engineers and managers to help design and deploy these new products. A successful candidate will have an established background in developing Machine Learning products with direct customer-facing experience, a strong technical ability and a motivation to achieve results. Experience in Machine Learning and software development is also a must. Responsible for the full development life cycle of our integrations and extensions for training support in Pytorch, JAX, and distributed training libraries with a focus on performance of latest ML models at scale on Trainium using latest techniques in performance optimization, accuracy, and resilience. you will lead the way to ensure support for key ML functionality in a combined chip / software platform, which will ensure the right thing is being built and delivered to customers Key job responsibilities Lead a team of engineers focused on enabling new ML training customers on the Neuron SDK / Trainium platform. Own the customer onboarding journey from model evaluation through production training at scale Drive engineering initiatives to maximize Model FLOPS Utilization (MFU) for customer workloads through performance analysis, profiling, and tuning tools. Build and maintain tooling, automation, and documentation that accelerates time-to-first-training for new customer models. Partner with compiler, runtime, and framework teams to identify and resolve blockers in customer workloads. Develop scalable processes for distributed training enablement (FSDP, DeepSpeed, Megatron, custom parallelism strategies). Drive technical strategy for supporting frontier model architectures (LLMs, MoE, multi-modal) on Trainium. Build strong cross-functional partnerships with product management, developer relations, and customer-facing teams. Recruit, mentor, and grow a high-performing team of ML systems engineers. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team About Us Inclusive Team Culture Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - Experience working with PyTorch or JAX software - 3+ years of engineering team management experience - 7+ years of working directly within engineering teams experience - 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - 3+ years of experience in deep learning / machine learning, including model training workflows - Experience with distributed training at scale (multi-node, multi-accelerator) PREFERRED QUALIFICATIONS - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience directly managing scientists or machine learning engineers - Experience debugging, profiling, and implementing best software engineering practices in large-scale systems - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels - Experience with performance analysis, profiling, and optimization for deep learning training workloads Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 212 700.00 USD annually
08/17/2026
Full time
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trainium-based servers that use them. As the SDM of Software Development for the Neuron Training team, you will be responsible for leading a strong team of engineers and managers to help design and deploy these new products. A successful candidate will have an established background in developing Machine Learning products with direct customer-facing experience, a strong technical ability and a motivation to achieve results. Experience in Machine Learning and software development is also a must. Responsible for the full development life cycle of our integrations and extensions for training support in Pytorch, JAX, and distributed training libraries with a focus on performance of latest ML models at scale on Trainium using latest techniques in performance optimization, accuracy, and resilience. you will lead the way to ensure support for key ML functionality in a combined chip / software platform, which will ensure the right thing is being built and delivered to customers Key job responsibilities Lead a team of engineers focused on enabling new ML training customers on the Neuron SDK / Trainium platform. Own the customer onboarding journey from model evaluation through production training at scale Drive engineering initiatives to maximize Model FLOPS Utilization (MFU) for customer workloads through performance analysis, profiling, and tuning tools. Build and maintain tooling, automation, and documentation that accelerates time-to-first-training for new customer models. Partner with compiler, runtime, and framework teams to identify and resolve blockers in customer workloads. Develop scalable processes for distributed training enablement (FSDP, DeepSpeed, Megatron, custom parallelism strategies). Drive technical strategy for supporting frontier model architectures (LLMs, MoE, multi-modal) on Trainium. Build strong cross-functional partnerships with product management, developer relations, and customer-facing teams. Recruit, mentor, and grow a high-performing team of ML systems engineers. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team About Us Inclusive Team Culture Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - Experience working with PyTorch or JAX software - 3+ years of engineering team management experience - 7+ years of working directly within engineering teams experience - 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - 3+ years of experience in deep learning / machine learning, including model training workflows - Experience with distributed training at scale (multi-node, multi-accelerator) PREFERRED QUALIFICATIONS - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience directly managing scientists or machine learning engineers - Experience debugging, profiling, and implementing best software engineering practices in large-scale systems - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels - Experience with performance analysis, profiling, and optimization for deep learning training workloads Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 212 700.00 USD annually
Sr Software Development Engineer, Neuron Collectives, Annapurna Labs
Annapurna Labs (U.S.) Inc. Cupertino, California
Annapurna Labs is an integral part of AWS and develops hardware and software components that are critical building blocks for EC2 infrastructure. We specialize in designing software, systems and chips that optimize the AWS customer experience. The AWS Neuron Collectives team is seeking a Software Engineer to optimize collective operations for AWS Trainium. Trainium is one of Amazon's highest priority initiatives, powering the frontier AI models being trained today. Collectives are the critical operations that scale AI compute across the data center. You'll work in depth to optimize compute for the specific topologies used to train modern LLMs. Working closely with the hardware team, you'll push for maximum performance using C/C++, interfacing with DMA and firmware and investigating detailed topologies. You'll analyze current collective algorithms using publicly accessible tools like Neuron Explorer and optimize these to fully utilize compute and bus bandwidth to scale across the data center. This is a unique opportunity to impact how AI training runs at AWS scale, while growing your technical breadth and depth. Key job responsibilities As a Neuron Collectives Software Developer, you will: Enhance collective algorithms and topologies for optimal training performance Use tools like Neuron Explorer to identify bottlenecks in compute and bus bandwidth utilization Monitor and analyze processor, DMA, firmware, and workload metrics Optimize collective operations to scale AI compute across the data center Work closely with the hardware team to co-optimize software and Trainium silicon Develop and optimize C/C++ implementations of collective communication patterns Investigate and implement improvements for specific training topologies used by modern LLMs Build and maintain analysis frameworks and automation solutions The role offers opportunities to work on cutting-edge AI training hardware while contributing to one of Amazon's most critical initiatives. A day in the life Inclusive Team Culture Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. About the team Annapurna Labs, part of AWS, created Trainium as a purpose-built AI training chip to revolutionize machine learning at Amazon scale. The Neuron Collectives team owns the software stack that enables collective operations - the communication primitives that allow AI training to scale across thousands of chips in the data center. Our work is essential to training the frontier models that power AI today. We work closely with hardware teams to extract maximum performance from Trainium, ensuring that compute and interconnect bandwidth are fully utilized. Our team sits at the intersection of hardware, firmware, and distributed systems. BASIC QUALIFICATIONS - Bachelor's degree in computer science or equivalent - 5+ years of Experience building complex software systems that have been successfully delivered to customers - 5+ years of Experience contributing to the architecture and design (architecture, design patterns, reliability and scaling) of new and current systems PREFERRED QUALIFICATIONS - Master's degree in computer science or equivalent - Familiarity with collective communication algorithms (e.g., all-reduce, all-gather) or distributed training frameworks Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 193 500.00 USD annually
08/17/2026
Full time
Annapurna Labs is an integral part of AWS and develops hardware and software components that are critical building blocks for EC2 infrastructure. We specialize in designing software, systems and chips that optimize the AWS customer experience. The AWS Neuron Collectives team is seeking a Software Engineer to optimize collective operations for AWS Trainium. Trainium is one of Amazon's highest priority initiatives, powering the frontier AI models being trained today. Collectives are the critical operations that scale AI compute across the data center. You'll work in depth to optimize compute for the specific topologies used to train modern LLMs. Working closely with the hardware team, you'll push for maximum performance using C/C++, interfacing with DMA and firmware and investigating detailed topologies. You'll analyze current collective algorithms using publicly accessible tools like Neuron Explorer and optimize these to fully utilize compute and bus bandwidth to scale across the data center. This is a unique opportunity to impact how AI training runs at AWS scale, while growing your technical breadth and depth. Key job responsibilities As a Neuron Collectives Software Developer, you will: Enhance collective algorithms and topologies for optimal training performance Use tools like Neuron Explorer to identify bottlenecks in compute and bus bandwidth utilization Monitor and analyze processor, DMA, firmware, and workload metrics Optimize collective operations to scale AI compute across the data center Work closely with the hardware team to co-optimize software and Trainium silicon Develop and optimize C/C++ implementations of collective communication patterns Investigate and implement improvements for specific training topologies used by modern LLMs Build and maintain analysis frameworks and automation solutions The role offers opportunities to work on cutting-edge AI training hardware while contributing to one of Amazon's most critical initiatives. A day in the life Inclusive Team Culture Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. About the team Annapurna Labs, part of AWS, created Trainium as a purpose-built AI training chip to revolutionize machine learning at Amazon scale. The Neuron Collectives team owns the software stack that enables collective operations - the communication primitives that allow AI training to scale across thousands of chips in the data center. Our work is essential to training the frontier models that power AI today. We work closely with hardware teams to extract maximum performance from Trainium, ensuring that compute and interconnect bandwidth are fully utilized. Our team sits at the intersection of hardware, firmware, and distributed systems. BASIC QUALIFICATIONS - Bachelor's degree in computer science or equivalent - 5+ years of Experience building complex software systems that have been successfully delivered to customers - 5+ years of Experience contributing to the architecture and design (architecture, design patterns, reliability and scaling) of new and current systems PREFERRED QUALIFICATIONS - Master's degree in computer science or equivalent - Familiarity with collective communication algorithms (e.g., all-reduce, all-gather) or distributed training frameworks Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 193 500.00 USD annually
Sr. Machine Learning - Compiler Engineer III, AWS Neuron, Annapurna Labs
Annapurna Labs (U.S.) Inc. Cupertino, California
Do you want to be part of AI revolution? At AWS our vision is to make deep learning pervasive for everyday developers and to democratize access to cutting-edge infrastructure. In order to deliver on that vision, we've created innovative software and hardware solutions that make it possible. AWS Neuron is the SDK that optimizes the performance of complex ML models executed on AWS Inferentia and Trainium, our custom chips designed to accelerate deep-learning workloads This role is for a senior software engineer in the Compiler team for AWS Neuron. As part of this role, you will be responsible for building next generation Neuron compiler which transforms ML models written in ML frameworks (e.g, PyTorch, TensorFlow, and JAX) to be deployed AWS Inferentia and Trainium based servers in the Amazon cloud. You will be responsible for solving hard compiler optimization problems to achieve optimum performance for variety of ML model families including massive scale large language models like Llama, Deepseek, and beyond as well as stable diffusion, vision transformers and multi-model models. You will be required to understand how these models work inside-out to make informed decisions on how to best coax the compiler to generate optimal implementation instruction. You will leverage your technical communications skill to partner with other teams and will be involved in pre-silicon design, bringing new products/features to market, and many other exciting projects. Experience in object-oriented languages like C++/Java is a must, experience with compilers or building ML models using ML frameworks on accelerators (e.g., GPUs) is preferred but not required. Experience with technologies like OpenXLA, StableHLO, MLIR will be added bonus! Explore the product and our history! AWS Utility Computing (UC) provides product innovations - from foundational services such as Amazon's Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS's services and features apart in the industry. As a member of the UC organization, you'll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. Key job responsibilities You will design, implement, test, deploy and maintain innovative software solutions to transform Neuron compiler's performance, stability and user-interface. You will work side by side with chip architects, runtime/OS engineers, scientists and ML Apps teams to seamlessly deploy cutting edge ML models from our customers on AWS accelerators with optimal cost/performance benefits. You will have opportunity to become front-face of Neuron Compiler to work with open-source communities (e.g., StableHLO, OpenXLA, MLIR) and influence industry wide partners to pioneer optimizing cutting-edge ML workloads on AWS software and hardware. You will also work on building innovative features that will deliver best possible experiences for our customers - developers across the globe. A day in the life As you design and code solutions to help our team drive efficiencies in compiler architecture, you'll create compiler optimization and verification passes, build features surface features and peculiarities of AWS accelerators to developers, implement tools to analyze numerical errors, and resolve the root cause of compiler defects. You'll also participate in design discussions, code review, and communicate with internal (other Neuron SDK and Amazon wide teams) and external stakeholders (open-source communities and respond to Neuron compiler related questions in open forums, e.g. GitHub). Lastly, work in a startup-like development environment, where you're always working on the most important stuff. About the team About the Team Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future. Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying. About AWS Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud. Mentorship & Career Growth We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. BASIC QUALIFICATIONS - 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience - 2+ years of experience in developing compiler features and optimizations - Proficiency with 1 or more of the following programming languages: C++ (preferred), C, Python PREFERRED QUALIFICATIONS - Master or PhD degree in computer science or equivalent - Proficiency with resource management, scheduling, code generation, and compute graph optimization - Experience optimizing Tensorflow, PyTorch or JAX deep learning models - Experience with multiple toolchains and Instruction Set Architectures Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 193 500.00 USD annually
08/17/2026
Full time
Do you want to be part of AI revolution? At AWS our vision is to make deep learning pervasive for everyday developers and to democratize access to cutting-edge infrastructure. In order to deliver on that vision, we've created innovative software and hardware solutions that make it possible. AWS Neuron is the SDK that optimizes the performance of complex ML models executed on AWS Inferentia and Trainium, our custom chips designed to accelerate deep-learning workloads This role is for a senior software engineer in the Compiler team for AWS Neuron. As part of this role, you will be responsible for building next generation Neuron compiler which transforms ML models written in ML frameworks (e.g, PyTorch, TensorFlow, and JAX) to be deployed AWS Inferentia and Trainium based servers in the Amazon cloud. You will be responsible for solving hard compiler optimization problems to achieve optimum performance for variety of ML model families including massive scale large language models like Llama, Deepseek, and beyond as well as stable diffusion, vision transformers and multi-model models. You will be required to understand how these models work inside-out to make informed decisions on how to best coax the compiler to generate optimal implementation instruction. You will leverage your technical communications skill to partner with other teams and will be involved in pre-silicon design, bringing new products/features to market, and many other exciting projects. Experience in object-oriented languages like C++/Java is a must, experience with compilers or building ML models using ML frameworks on accelerators (e.g., GPUs) is preferred but not required. Experience with technologies like OpenXLA, StableHLO, MLIR will be added bonus! Explore the product and our history! AWS Utility Computing (UC) provides product innovations - from foundational services such as Amazon's Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS's services and features apart in the industry. As a member of the UC organization, you'll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. Key job responsibilities You will design, implement, test, deploy and maintain innovative software solutions to transform Neuron compiler's performance, stability and user-interface. You will work side by side with chip architects, runtime/OS engineers, scientists and ML Apps teams to seamlessly deploy cutting edge ML models from our customers on AWS accelerators with optimal cost/performance benefits. You will have opportunity to become front-face of Neuron Compiler to work with open-source communities (e.g., StableHLO, OpenXLA, MLIR) and influence industry wide partners to pioneer optimizing cutting-edge ML workloads on AWS software and hardware. You will also work on building innovative features that will deliver best possible experiences for our customers - developers across the globe. A day in the life As you design and code solutions to help our team drive efficiencies in compiler architecture, you'll create compiler optimization and verification passes, build features surface features and peculiarities of AWS accelerators to developers, implement tools to analyze numerical errors, and resolve the root cause of compiler defects. You'll also participate in design discussions, code review, and communicate with internal (other Neuron SDK and Amazon wide teams) and external stakeholders (open-source communities and respond to Neuron compiler related questions in open forums, e.g. GitHub). Lastly, work in a startup-like development environment, where you're always working on the most important stuff. About the team About the Team Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future. Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying. About AWS Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud. Mentorship & Career Growth We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. BASIC QUALIFICATIONS - 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience - 2+ years of experience in developing compiler features and optimizations - Proficiency with 1 or more of the following programming languages: C++ (preferred), C, Python PREFERRED QUALIFICATIONS - Master or PhD degree in computer science or equivalent - Proficiency with resource management, scheduling, code generation, and compute graph optimization - Experience optimizing Tensorflow, PyTorch or JAX deep learning models - Experience with multiple toolchains and Instruction Set Architectures Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 193 500.00 USD annually
Neuron Runtime Software Development Engineer , Neuron Runtime
Annapurna Labs (U.S.) Inc. Cupertino, California
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the servers that use them. As the Software Development Engineer for the Neuron Runtime Team, you will be responsible for working alongside a team of engineers to develop and maintain high-performance runtime libraries and drivers for machine learning applications and AI accelerators. You will work on design, development, and deployment of Neuron Runtime and other Neuron components. The profiler plays a crucial role to internal and external customers in optimizing AI workloads across hardware platforms such as Trainium and Inferentia devices, by providing deep insights into performance bottlenecks and system behavior. Improving performance of ML Kernels and ML Frameworks. In this role, you will manage the full development life cycle of the Neuron Runtime, ensuring scalability, reliability, and usability. You will collaborate with cross-functional teams to ensure that the our C++ compiler generates key information so customers can understand and optimize the performance of our custom hardware. Additionally, you will drive innovations that allow the profiler to support multiple frameworks, such as PyTorch, JAX, and XLA. A successful candidate will have experience in architecting, building, and operating distributed systems with a focus on high availability and fault tolerance, Hands-on experience with AWS services (e.g., EC2, ECS, CloudWatch, S3, Lambda) in production environments and track record in Owning services end-to-end including deployment, monitoring, alarming, on-call, and post-incident review. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually USA, WA, Seattle - 143 400.00 USD annually
08/17/2026
Full time
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the servers that use them. As the Software Development Engineer for the Neuron Runtime Team, you will be responsible for working alongside a team of engineers to develop and maintain high-performance runtime libraries and drivers for machine learning applications and AI accelerators. You will work on design, development, and deployment of Neuron Runtime and other Neuron components. The profiler plays a crucial role to internal and external customers in optimizing AI workloads across hardware platforms such as Trainium and Inferentia devices, by providing deep insights into performance bottlenecks and system behavior. Improving performance of ML Kernels and ML Frameworks. In this role, you will manage the full development life cycle of the Neuron Runtime, ensuring scalability, reliability, and usability. You will collaborate with cross-functional teams to ensure that the our C++ compiler generates key information so customers can understand and optimize the performance of our custom hardware. Additionally, you will drive innovations that allow the profiler to support multiple frameworks, such as PyTorch, JAX, and XLA. A successful candidate will have experience in architecting, building, and operating distributed systems with a focus on high availability and fault tolerance, Hands-on experience with AWS services (e.g., EC2, ECS, CloudWatch, S3, Lambda) in production environments and track record in Owning services end-to-end including deployment, monitoring, alarming, on-call, and post-incident review. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions. This organization builds the full stack of software, servers and chips to accelerate at the highest scale. About the team Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually USA, WA, Seattle - 143 400.00 USD annually
Software Engineer Neuron Containers, Annapurna Labs, Neuron Containers, Annapurna Labs
Annapurna Labs (U.S.) Inc. Cupertino, California
Annapurna Labs was a startup acquired by AWS in 2015 and is now fully integrated. AWS Neuron is the complete software stack for the Inferentia and Trainium ML accelerators - the custom silicon that powers large-scale AI workloads on AWS. The Neuron Containers team is looking for a Software Development Engineer to build platform integrations that enable customers to run distributed training and inference workloads on Neuron at scale. The team owns Neuron integration with Kubernetes, ECS, and Slurm - handling device allocation, fault tolerance, auto-scaling, and orchestration across large clusters. The team also owns delivery of Neuron Deep Learning Containers (DLCs) and Deep Learning AMIs (DLAMIs) - the pre-configured container images and machine images that package the Neuron SDK for customer deployment on Trainium and Inferentia instances. Key job responsibilities Design and implement container platform integrations - device plugins, DRA drivers, and operators for ML accelerator resource management Build and maintain Neuron DLCs and DLAMIs for customer deployment across EKS, ECS, EC2, and SageMaker Diagnose and resolve performance and scalability issues across large customer clusters Simplify systems - deprecate legacy software and reduce complexity in container delivery pipelines Deliver software across the full development lifecycle including design documentation, implementation, testing, deployment, and operations A day in the life You'll work with teams across the Neuron org and customers to build and maintain integrations for current and next-generation accelerators. You'll participate in architecture reviews, triage test failures, resolve operational issues, and contribute to upstream Kubernetes projects. You'll debug platform integration problems - how Neuron interacts with container runtimes, orchestrators, and scheduling systems a scale. BASIC QUALIFICATIONS - 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent - Experience with distributed systems or large-scale cluster infrastructure - Familiarity with ML training/inference workflows (distributed training, collective - Experience with AWS compute services (EC2, EKS, ECS, ECR) - Hands-on experience with Helm, Prometheus, or Kubernetes operator frameworks - Experience with container image pipelines, Deep Learning Containers, or Deep Learning AMIs - Contributions to open-source projects, particularly in the Kubernetes ecosystem Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually
08/17/2026
Full time
Annapurna Labs was a startup acquired by AWS in 2015 and is now fully integrated. AWS Neuron is the complete software stack for the Inferentia and Trainium ML accelerators - the custom silicon that powers large-scale AI workloads on AWS. The Neuron Containers team is looking for a Software Development Engineer to build platform integrations that enable customers to run distributed training and inference workloads on Neuron at scale. The team owns Neuron integration with Kubernetes, ECS, and Slurm - handling device allocation, fault tolerance, auto-scaling, and orchestration across large clusters. The team also owns delivery of Neuron Deep Learning Containers (DLCs) and Deep Learning AMIs (DLAMIs) - the pre-configured container images and machine images that package the Neuron SDK for customer deployment on Trainium and Inferentia instances. Key job responsibilities Design and implement container platform integrations - device plugins, DRA drivers, and operators for ML accelerator resource management Build and maintain Neuron DLCs and DLAMIs for customer deployment across EKS, ECS, EC2, and SageMaker Diagnose and resolve performance and scalability issues across large customer clusters Simplify systems - deprecate legacy software and reduce complexity in container delivery pipelines Deliver software across the full development lifecycle including design documentation, implementation, testing, deployment, and operations A day in the life You'll work with teams across the Neuron org and customers to build and maintain integrations for current and next-generation accelerators. You'll participate in architecture reviews, triage test failures, resolve operational issues, and contribute to upstream Kubernetes projects. You'll debug platform integration problems - how Neuron interacts with container runtimes, orchestrators, and scheduling systems a scale. BASIC QUALIFICATIONS - 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language PREFERRED QUALIFICATIONS - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent - Experience with distributed systems or large-scale cluster infrastructure - Familiarity with ML training/inference workflows (distributed training, collective - Experience with AWS compute services (EC2, EKS, ECS, ECR) - Hands-on experience with Helm, Prometheus, or Kubernetes operator frameworks - Experience with container image pipelines, Deep Learning Containers, or Deep Learning AMIs - Contributions to open-source projects, particularly in the Kubernetes ecosystem Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually
Software Development Manager - Compiler
Annapurna Labs (U.S.) Inc. Cupertino, California
The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. The Inferentia chip delivers best-in-class ML inference performance at the lowest cost in cloud. Trainium delivers the best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, runtime and natively integrates into popular ML frameworks, such as PyTorch and JAX. .AWS Neuron and Trainium are used at scale with customers and partners like PyTorch, Anthropic, Poolside, Decart, Epic Games, Snap, AirBnB, Autodesk, Amazon Alexa, and more customers in various other segments. The Team: The Amazon Annapurna Labs team is a responsible for building innovation in silicon and software for AWS customers. We are at the forefront of innovation by combining cloud scale with the world's most talented engineers. Our team covers multiple disciplines including silicon engineering, hardware design and verification, software and operations. With such breadth of talent, there's opportunity to learn all of the time. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. When you couple that with the ability to work on so many different products and services, it's a very unique learning culture. Learn more about Our History: You: We are seeking a talented SW Engineering Manager with strong leadership/ mentoring skills to join our Deep Learning Compiler Team. As a Manager III you be leading a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You'll need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron Manager, innovating on behalf of our customers. You'll leverage your technical communications skill as a hands-on partner to AWS ML services teams, involved in pre-silicon design, bringing new products/features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that will boost the entire deep learning community. You will have deep knowledge of resource management, scheduling, code generation, optimization, and new instruction architectures including CPU, NPU, GPU and novel forms of compute. Explore the Product: In order to be considered for this role, candidates must be currently located or willing to relocate to Cupertino (preferred), Seattle, or Austin. About the team Inclusive Team Culture Here at Annapurna Labs, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - 5+ years of engineering team management experience - 9+ years of working directly within engineering teams experience - 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - Deep understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization) - Strong software design fundamentals and excellent system-level coding skills with an emphasis on graph theory and performance techniques PREFERRED QUALIFICATIONS - PhD in computer science, computer engineering, or related field, or MS degree - Experience with general troubleshooting/debugging of hardware, or experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience with XLA, TVM, MLIR, LLVM, deep learning models and algorithms, and deep learning framework design. - Interactions with open-source communities, in either a leadership or code contributor role Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 212 700.00 USD annually
08/17/2026
Full time
The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. The Inferentia chip delivers best-in-class ML inference performance at the lowest cost in cloud. Trainium delivers the best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, runtime and natively integrates into popular ML frameworks, such as PyTorch and JAX. .AWS Neuron and Trainium are used at scale with customers and partners like PyTorch, Anthropic, Poolside, Decart, Epic Games, Snap, AirBnB, Autodesk, Amazon Alexa, and more customers in various other segments. The Team: The Amazon Annapurna Labs team is a responsible for building innovation in silicon and software for AWS customers. We are at the forefront of innovation by combining cloud scale with the world's most talented engineers. Our team covers multiple disciplines including silicon engineering, hardware design and verification, software and operations. With such breadth of talent, there's opportunity to learn all of the time. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. When you couple that with the ability to work on so many different products and services, it's a very unique learning culture. Learn more about Our History: You: We are seeking a talented SW Engineering Manager with strong leadership/ mentoring skills to join our Deep Learning Compiler Team. As a Manager III you be leading a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You'll need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron Manager, innovating on behalf of our customers. You'll leverage your technical communications skill as a hands-on partner to AWS ML services teams, involved in pre-silicon design, bringing new products/features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that will boost the entire deep learning community. You will have deep knowledge of resource management, scheduling, code generation, optimization, and new instruction architectures including CPU, NPU, GPU and novel forms of compute. Explore the Product: In order to be considered for this role, candidates must be currently located or willing to relocate to Cupertino (preferred), Seattle, or Austin. About the team Inclusive Team Culture Here at Annapurna Labs, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - 5+ years of engineering team management experience - 9+ years of working directly within engineering teams experience - 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - Deep understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization) - Strong software design fundamentals and excellent system-level coding skills with an emphasis on graph theory and performance techniques PREFERRED QUALIFICATIONS - PhD in computer science, computer engineering, or related field, or MS degree - Experience with general troubleshooting/debugging of hardware, or experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience with XLA, TVM, MLIR, LLVM, deep learning models and algorithms, and deep learning framework design. - Interactions with open-source communities, in either a leadership or code contributor role Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 212 700.00 USD annually
Software Development Manager - NKI Compiler, AWS Neuron, Annapurna Labs
Annapurna Labs (U.S.) Inc. Cupertino, California
The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. Trainium delivers best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, the Neuron Kernel Interface (NKI) compiler, and a runtime that natively integrates into popular ML frameworks such as PyTorch and JAX. Neuron Kernel Interface (NKI) is a bare-metal language and compiler for directly programming NeuronDevices available on AWS Trainium instances. You can use NKI to develop, optimize, and run new operators directly on NeuronCores while making full use of available compute and memory resources. Explore NKI: - AWS Neuron is used at scale by customers such as Epic Games, Snap, Airbnb, Autodesk, Amazon Alexa, and Amazon Rekognition, along with many others across a range of segments. The Team: The Amazon Annapurna Labs team is responsible for building innovative silicon and software for AWS customers. We are at the forefront of innovation, combining cloud scale with the world's most talented engineers. Our team covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. With such breadth of talent, there is opportunity to learn all of the time. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. When you couple that with the ability to work on so many different products and services, it makes for a unique learning culture. Learn more about our history: You: We are seeking a talented Software Development Manager with strong leadership and mentoring skills to join our NKI development team. As an SDM III, you will lead a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You will need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron manager, innovating on behalf of our customers. You will draw on knowledge of resource management, scheduling, code generation, optimization, and instruction architectures across CPU, NPU, GPU, and novel forms of compute. You will leverage your technical communication skills to partner with AWS ML services teams and pre-silicon design, and to bring new products and features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that boosts the entire deep learning community. Explore the Product: - - In order to be considered for this role, candidates must be currently located in or willing to relocate to Cupertino, CA. Key job responsibilities - Lead, grow, and mentor a team of compiler engineers, including hiring, career development, and team health - Own execution and delivery of NKI compiler features across release cycles, balancing scope, quality, and timelines - Set technical direction and roadmap for your area of the NKI compiler in partnership with senior engineers - Partner across teams, including frameworks, kernel development, runtime, hardware/pre-silicon design, and product management, to bring new features to market - Drive resolution of complex, high-priority technical issues, staying close enough to the code and hardware to guide the team - Represent your team's work and priorities to senior leadership and cross-organizational stakeholders - Anchor decisions in customer impact, ensuring the team builds what unblocks real ML workloads on Trainium A day in the life No two days look the same on the NKI team, but most blend people leadership, technical depth, and cross-team collaboration. You might start the morning in a 1:1 with an engineer, working through the design of a new compiler optimization or unblocking a tricky scheduling problem. Mid-morning, you join a release sync to check in on branch-cut readiness and make a call on what makes the current train versus the next. After lunch, you catch up with a partner team, kernel developers, runtime, or hardware design, to align on an upcoming feature and its dependencies. In the afternoon, you spend focused time on the roadmap: shaping where NKI is heading over the next few quarters, and translating customer needs into concrete technical investments. You review a design doc, leave feedback that sharpens the team's thinking, and dig into a profiling result yourself to understand where a real workload is leaving performance on the table. You close the day by clearing a path for your team and resolving an escalation, connecting two people who should be talking, or writing up a decision so the team can move fast tomorrow. Throughout, you keep one question at the center: what does this unblock for our customers? You are technically credible enough to earn your team's trust, and you spend your energy on the highest-leverage problems: growing your people, delivering the compiler, and inventing on behalf of the customers who run their most demanding ML workloads on Trainium. About the team Inclusive Team Culture Here at Annapurna Labs, we embrace our differences. We are committed to furthering our culture of inclusion. Amazon has ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - 2+ years of engineering team management experience - 6+ years of working directly within engineering teams experience - 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - Understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization) - Strong software design fundamentals and excellent system-level coding skills PREFERRED QUALIFICATIONS - M.S. or Ph.D. in Computer Science or related technical field Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support . click apply for full job details
08/17/2026
Full time
The Product: AWS Machine Learning accelerators are at the forefront of AWS innovation. Trainium delivers best-in-class ML training performance with the most teraflops (TFLOPS) of compute power for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler, the Neuron Kernel Interface (NKI) compiler, and a runtime that natively integrates into popular ML frameworks such as PyTorch and JAX. Neuron Kernel Interface (NKI) is a bare-metal language and compiler for directly programming NeuronDevices available on AWS Trainium instances. You can use NKI to develop, optimize, and run new operators directly on NeuronCores while making full use of available compute and memory resources. Explore NKI: - AWS Neuron is used at scale by customers such as Epic Games, Snap, Airbnb, Autodesk, Amazon Alexa, and Amazon Rekognition, along with many others across a range of segments. The Team: The Amazon Annapurna Labs team is responsible for building innovative silicon and software for AWS customers. We are at the forefront of innovation, combining cloud scale with the world's most talented engineers. Our team covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. With such breadth of talent, there is opportunity to learn all of the time. We operate in spaces that are very large, yet our teams remain small and agile. There is no blueprint. We're inventing. We're experimenting. When you couple that with the ability to work on so many different products and services, it makes for a unique learning culture. Learn more about our history: You: We are seeking a talented Software Development Manager with strong leadership and mentoring skills to join our NKI development team. As an SDM III, you will lead a team of experienced compiler engineers developing compiler optimization algorithms and deploying, at scale, a new compiler targeting AWS custom hardware. You will need to be technically capable, credible, and curious in your own right as a trusted AWS Neuron manager, innovating on behalf of our customers. You will draw on knowledge of resource management, scheduling, code generation, optimization, and instruction architectures across CPU, NPU, GPU, and novel forms of compute. You will leverage your technical communication skills to partner with AWS ML services teams and pre-silicon design, and to bring new products and features to market. As deep learning models become more versatile, using compiler technologies to achieve both high performance and high productivity becomes essential. Join the team to build the software that boosts the entire deep learning community. Explore the Product: - - In order to be considered for this role, candidates must be currently located in or willing to relocate to Cupertino, CA. Key job responsibilities - Lead, grow, and mentor a team of compiler engineers, including hiring, career development, and team health - Own execution and delivery of NKI compiler features across release cycles, balancing scope, quality, and timelines - Set technical direction and roadmap for your area of the NKI compiler in partnership with senior engineers - Partner across teams, including frameworks, kernel development, runtime, hardware/pre-silicon design, and product management, to bring new features to market - Drive resolution of complex, high-priority technical issues, staying close enough to the code and hardware to guide the team - Represent your team's work and priorities to senior leadership and cross-organizational stakeholders - Anchor decisions in customer impact, ensuring the team builds what unblocks real ML workloads on Trainium A day in the life No two days look the same on the NKI team, but most blend people leadership, technical depth, and cross-team collaboration. You might start the morning in a 1:1 with an engineer, working through the design of a new compiler optimization or unblocking a tricky scheduling problem. Mid-morning, you join a release sync to check in on branch-cut readiness and make a call on what makes the current train versus the next. After lunch, you catch up with a partner team, kernel developers, runtime, or hardware design, to align on an upcoming feature and its dependencies. In the afternoon, you spend focused time on the roadmap: shaping where NKI is heading over the next few quarters, and translating customer needs into concrete technical investments. You review a design doc, leave feedback that sharpens the team's thinking, and dig into a profiling result yourself to understand where a real workload is leaving performance on the table. You close the day by clearing a path for your team and resolving an escalation, connecting two people who should be talking, or writing up a decision so the team can move fast tomorrow. Throughout, you keep one question at the center: what does this unblock for our customers? You are technically credible enough to earn your team's trust, and you spend your energy on the highest-leverage problems: growing your people, delivering the compiler, and inventing on behalf of the customers who run their most demanding ML workloads on Trainium. About the team Inclusive Team Culture Here at Annapurna Labs, we embrace our differences. We are committed to furthering our culture of inclusion. Amazon has ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. BASIC QUALIFICATIONS - 2+ years of engineering team management experience - 6+ years of working directly within engineering teams experience - 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Experience partnering with product or program management teams - Understanding of compilers (resource management, instruction scheduling, code generation, and compute graph optimization) - Strong software design fundamentals and excellent system-level coding skills PREFERRED QUALIFICATIONS - M.S. or Ph.D. in Computer Science or related technical field Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support . click apply for full job details
  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board