it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Register as Employer
  • Contact us
  • Career Advice
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

20 jobs found

Email me jobs like this
Refine Search
Current Search
hpc gpu systems engineer
Platform Engineer, Model Shaping
Together AI Remote, Oregon
About the Role The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition to that, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems. As a Platform Engineer in Model Shaping, you will work at the intersection of backend engineering and infrastructure, building the foundational layers of Together's platform for model customization and evaluation. You will design, develop, and operate both the backend services and the underlying systems that enable us to sustainably and reliably scale production workflows launched by our users, as well as internal research experiments. You will operate in a cross-functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on. You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams. Responsibilities Design and build Together's systems and infrastructure for model customization, including user-facing features and internal improvements Contribute to reliability improvements for the platform, participating in an on-call rotation and improving processes for incident response Create and improve internal tooling for deployment, continuous integration, and observability Build a job orchestration platform spanning multiple datacenters, supporting a highly heterogeneous hardware landscape Partner with teams developing internal services, co-designing these services and incorporating them in systems built within Together Requirements 3+ years of experience in building infrastructure or backend components of production services Extensive experience designing, operating, and troubleshooting production Linux environments and Kubernetes-based platforms Strong software engineering background in Python or Go Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD) Cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare-metal/cloud environment Strong communication skills, be willing to document systems and processes and collaborate with peers of varying technical expertise Comfortable operating across the stack, from cluster operations and infrastructure automation to backend service development Experience in any of the following will make you stand out: Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters, ideally with hands-on experience in operating NVIDIA's networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA) Deployment of services for AI training or inference Networking fundamentals, including TCP/IP, DNS, routing, load balancing, TLS, and network debugging tools Maintaining or contributing to open-source projects About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
09/25/2026
Full time
About the Role The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition to that, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems. As a Platform Engineer in Model Shaping, you will work at the intersection of backend engineering and infrastructure, building the foundational layers of Together's platform for model customization and evaluation. You will design, develop, and operate both the backend services and the underlying systems that enable us to sustainably and reliably scale production workflows launched by our users, as well as internal research experiments. You will operate in a cross-functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on. You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams. Responsibilities Design and build Together's systems and infrastructure for model customization, including user-facing features and internal improvements Contribute to reliability improvements for the platform, participating in an on-call rotation and improving processes for incident response Create and improve internal tooling for deployment, continuous integration, and observability Build a job orchestration platform spanning multiple datacenters, supporting a highly heterogeneous hardware landscape Partner with teams developing internal services, co-designing these services and incorporating them in systems built within Together Requirements 3+ years of experience in building infrastructure or backend components of production services Extensive experience designing, operating, and troubleshooting production Linux environments and Kubernetes-based platforms Strong software engineering background in Python or Go Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD) Cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare-metal/cloud environment Strong communication skills, be willing to document systems and processes and collaborate with peers of varying technical expertise Comfortable operating across the stack, from cluster operations and infrastructure automation to backend service development Experience in any of the following will make you stand out: Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters, ideally with hands-on experience in operating NVIDIA's networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA) Deployment of services for AI training or inference Networking fundamentals, including TCP/IP, DNS, routing, load balancing, TLS, and network debugging tools Maintaining or contributing to open-source projects About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
Nvidia
Senior Software Engineer, NCCL
Nvidia Santa Clara, California
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning. What you will be doing: Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters. Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components. Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware. What we need to see: M.S./Ph.D. degree in CS/CE or equivalent experience. 5+ years of relevant experience. Excellent C/C++ programming and debugging skills. Strong experience with Linux. Expert understanding of computer system architecture and operating systems. Experience with parallel programming interfaces and communication runtimes. Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment. Ways to stand out from the crowd: Deep understanding of technology and passionate about what you do. Experience with CUDA programming and NVIDIA GPUs. Knowledge of high-performance networks like InfiniBand, iWARP etc. Experience with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, etc. Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment. NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning. What you will be doing: Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters. Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components. Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware. What we need to see: M.S./Ph.D. degree in CS/CE or equivalent experience. 5+ years of relevant experience. Excellent C/C++ programming and debugging skills. Strong experience with Linux. Expert understanding of computer system architecture and operating systems. Experience with parallel programming interfaces and communication runtimes. Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment. Ways to stand out from the crowd: Deep understanding of technology and passionate about what you do. Experience with CUDA programming and NVIDIA GPUs. Knowledge of high-performance networks like InfiniBand, iWARP etc. Experience with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, etc. Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment. NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Deep Learning Framework Communications Engineer
Nvidia Westford, Massachusetts
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Deep Learning Framework Communications Engineer
Nvidia Santa Clara, California
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Deep Learning Framework Communications Engineer
Nvidia Austin, Texas
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Deep Learning Framework Communications Engineer
Nvidia Durham, North Carolina
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Linux Kernel Systems Software Engineer - CSP Engagements
Nvidia Austin, Texas
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Linux Kernel Systems Software Engineer - CSP Engagements
Nvidia Redmond, Washington
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Linux Kernel Systems Software Engineer - CSP Engagements
Nvidia Santa Clara, California
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Senior Linux Kernel Systems Software Engineer - CSP Engagements
Nvidia Seattle, Washington
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Customer Success Engineer (CSE), GPU Cluster
Together AI San Francisco, California
About the role As a Customer Success Engineer at Together AI, you will serve as the named technical owner for one of our most strategic customer relationships. You will be the primary technical point of contact across all infrastructure domains - compute, networking, storage, and facilities - ensuring flawless delivery and operational health of large-scale GPU deployments. This role sits at the intersection of deep infrastructure expertise and high-stakes customer partnership, making you a critical driver of both customer success and company growth. Responsibilities Serve as the named technical point of contact for a dedicated strategic customer, owning the end-to-end technical relationship across compute, networking, storage, and facilities Drive structured engagement through regular cadences - status reporting, technical steering meetings, quarterly business reviews (QBRs), and executive business reviews (EBRs) - spanning both operational and strategic levels Translate customer operational feedback into actionable input for Engineering, Product, and Infrastructure roadmaps Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains in partnership with Support, SRE, DC Ops, and Engineering teams Own end-to-end RMA coordination and hardware lifecycle management, including acceptance testing, spare inventory management, and hardware health reporting for large-scale GPU deployments Maintain deep technical expertise across the customer's infrastructure stack - GPU compute, high-speed fabric, and large-scale storage systems - advising on configuration, operational best practices, and incident resolution Own the observability strategy for the customer estate, including alert policy definition, dashboard development, and proactive health management across all infrastructure layers Coordinate DC operations and facilities events in partnership with internal teams and hosting providers, ensuring SLA compliance and cluster availability Act as project manager for all capacity expansions, owning the full node deployment lifecycle from freight receipt through production acceptance Qualifications 5+ years in a customer-facing technical role, with 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure Deep expertise in GPU infrastructure - GPU health diagnostics, RMA workflows, and hardware acceptance testing Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture Working knowledge of enterprise storage systems, including high-density NVMe, parallel file systems, and metadata infrastructure Experience with DC operations, facilities coordination, and hosting provider SLA management Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent) Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards Proficiency in Python, Bash, or infrastructure automation tools preferred About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $260-290K OTE + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Location San Francisco, CA (Hybrid) or New York, NY (Hybrid) Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our Privacy Policy at
09/25/2026
Full time
About the role As a Customer Success Engineer at Together AI, you will serve as the named technical owner for one of our most strategic customer relationships. You will be the primary technical point of contact across all infrastructure domains - compute, networking, storage, and facilities - ensuring flawless delivery and operational health of large-scale GPU deployments. This role sits at the intersection of deep infrastructure expertise and high-stakes customer partnership, making you a critical driver of both customer success and company growth. Responsibilities Serve as the named technical point of contact for a dedicated strategic customer, owning the end-to-end technical relationship across compute, networking, storage, and facilities Drive structured engagement through regular cadences - status reporting, technical steering meetings, quarterly business reviews (QBRs), and executive business reviews (EBRs) - spanning both operational and strategic levels Translate customer operational feedback into actionable input for Engineering, Product, and Infrastructure roadmaps Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains in partnership with Support, SRE, DC Ops, and Engineering teams Own end-to-end RMA coordination and hardware lifecycle management, including acceptance testing, spare inventory management, and hardware health reporting for large-scale GPU deployments Maintain deep technical expertise across the customer's infrastructure stack - GPU compute, high-speed fabric, and large-scale storage systems - advising on configuration, operational best practices, and incident resolution Own the observability strategy for the customer estate, including alert policy definition, dashboard development, and proactive health management across all infrastructure layers Coordinate DC operations and facilities events in partnership with internal teams and hosting providers, ensuring SLA compliance and cluster availability Act as project manager for all capacity expansions, owning the full node deployment lifecycle from freight receipt through production acceptance Qualifications 5+ years in a customer-facing technical role, with 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure Deep expertise in GPU infrastructure - GPU health diagnostics, RMA workflows, and hardware acceptance testing Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture Working knowledge of enterprise storage systems, including high-density NVMe, parallel file systems, and metadata infrastructure Experience with DC operations, facilities coordination, and hosting provider SLA management Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent) Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards Proficiency in Python, Bash, or infrastructure automation tools preferred About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $260-290K OTE + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Location San Francisco, CA (Hybrid) or New York, NY (Hybrid) Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our Privacy Policy at
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AI Remote, Oregon
About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world's largest AI training and inference workloads. You'll manage and scale high-performance parallel filesystems and object stores, evaluate and integrate cutting-edge technologies such as Vast, Weka, Ceph, and Lustre, and solve the complex engineering challenges of operating at extreme throughput, low-latency data paths, and massive cluster-scale storage operations. You will also build Kubernetes-native storage operators and self-service platforms that provide automated provisioning, strict multi-tenancy, performance isolation, and quota enforcement at cluster scale. Day-to-day, you'll optimize end-to-end data paths for 10-50 GB/s per node, design multi-tier caching architectures, implement intelligent prefetching and model-weight distribution, and tune parallel filesystems for AI workloads. Responsibilities Architect and implement the technical strategy and storage roadmap for Together AI, driving high-performance architectural decisions as we scale our GPU fleet. Engineer and scale multi-petabyte AI/ML storage systems by integrating Vast, Weka, and Ceph while executing deep cost optimization through automated tiering and lifecycle policies. Develop intelligent caching and tiered storage architectures to achieve extreme IOPS and cluster-wide throughput at GPU scale for training and inference workloads. Tune storage isolation at the L2/L3 network layers to ensure secure, production-grade multi-tenancy for storage clients. Code Kubernetes storage operators and controllers to enable automated provisioning, self-service abstractions, and quota enforcement. Engineer end-to-end data paths to achieve 10+ GB/s per GPU node; architect multi-tier caching for model weights and datasets; tune parallel filesystems using advanced profiling; and scale storage infrastructure across thousands of nodes. Optimize end-to-end data paths through advanced benchmarking and profiling, contributing high-impact code to open-source storage projects and internal tooling. Requirements 8+ years in storage engineering, managing distributed storage at multi-petabyte scale Proven track record deploying and operating high-performance storage for GPU/HPC clusters Deep Kubernetes and cloud-native storage experience in production environments Strong coding skills in Go and Python with demonstrated ability to build production-grade systems and tooling BS/MS in Computer Science, Engineering, or equivalent practical experience History of technical leadership: designing systems that significantly improved performance, reliability (99.999%+ uptime), or cost efficiency Distributed Storage Systems: Deep expertise in either of Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale Object Storage: Production experience with S3, MinIO, Ceph, or R2 including performance optimization and cost management Kubernetes Storage: CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers Storage optimization for GPU workloads, RDMA/InfiniBand networking, parallel filesystem optimization (TB/s aggregate cluster throughput - line saturation) Programming: Go and Python for automation, operators, and tooling Infrastructure as Code: Terraform, Ansible, Helm, GitOps (ArgoCD) Linux Storage Stack: Advanced knowledge of filesystems (ext4, xfs), LVM, NVMe optimization, RAID configurations Observability: Prometheus, Grafana, Thanos architecture and operations Nice to Have Skills GPU Direct Storage (GDS), NVMe-oF, storage networking, RDMA implementations ML/AI storage patterns (model weights, checkpointing, dataset caching) Storage benchmarking and profiling tools (fio, iperf3, iostat, blktrace). About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $250,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
09/25/2026
Full time
About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world's largest AI training and inference workloads. You'll manage and scale high-performance parallel filesystems and object stores, evaluate and integrate cutting-edge technologies such as Vast, Weka, Ceph, and Lustre, and solve the complex engineering challenges of operating at extreme throughput, low-latency data paths, and massive cluster-scale storage operations. You will also build Kubernetes-native storage operators and self-service platforms that provide automated provisioning, strict multi-tenancy, performance isolation, and quota enforcement at cluster scale. Day-to-day, you'll optimize end-to-end data paths for 10-50 GB/s per node, design multi-tier caching architectures, implement intelligent prefetching and model-weight distribution, and tune parallel filesystems for AI workloads. Responsibilities Architect and implement the technical strategy and storage roadmap for Together AI, driving high-performance architectural decisions as we scale our GPU fleet. Engineer and scale multi-petabyte AI/ML storage systems by integrating Vast, Weka, and Ceph while executing deep cost optimization through automated tiering and lifecycle policies. Develop intelligent caching and tiered storage architectures to achieve extreme IOPS and cluster-wide throughput at GPU scale for training and inference workloads. Tune storage isolation at the L2/L3 network layers to ensure secure, production-grade multi-tenancy for storage clients. Code Kubernetes storage operators and controllers to enable automated provisioning, self-service abstractions, and quota enforcement. Engineer end-to-end data paths to achieve 10+ GB/s per GPU node; architect multi-tier caching for model weights and datasets; tune parallel filesystems using advanced profiling; and scale storage infrastructure across thousands of nodes. Optimize end-to-end data paths through advanced benchmarking and profiling, contributing high-impact code to open-source storage projects and internal tooling. Requirements 8+ years in storage engineering, managing distributed storage at multi-petabyte scale Proven track record deploying and operating high-performance storage for GPU/HPC clusters Deep Kubernetes and cloud-native storage experience in production environments Strong coding skills in Go and Python with demonstrated ability to build production-grade systems and tooling BS/MS in Computer Science, Engineering, or equivalent practical experience History of technical leadership: designing systems that significantly improved performance, reliability (99.999%+ uptime), or cost efficiency Distributed Storage Systems: Deep expertise in either of Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale Object Storage: Production experience with S3, MinIO, Ceph, or R2 including performance optimization and cost management Kubernetes Storage: CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers Storage optimization for GPU workloads, RDMA/InfiniBand networking, parallel filesystem optimization (TB/s aggregate cluster throughput - line saturation) Programming: Go and Python for automation, operators, and tooling Infrastructure as Code: Terraform, Ansible, Helm, GitOps (ArgoCD) Linux Storage Stack: Advanced knowledge of filesystems (ext4, xfs), LVM, NVMe optimization, RAID configurations Observability: Prometheus, Grafana, Thanos architecture and operations Nice to Have Skills GPU Direct Storage (GDS), NVMe-oF, storage networking, RDMA implementations ML/AI storage patterns (model weights, checkpointing, dataset caching) Storage benchmarking and profiling tools (fio, iperf3, iostat, blktrace). About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $250,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
Nvidia
Principal Software Engineer, Rack-Scale System Software - CSP Engagements
Nvidia Austin, Texas
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Principal Software Engineer, Rack-Scale System Software - CSP Engagements
Nvidia Santa Clara, California
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Principal Software Engineer, Rack-Scale System Software - CSP Engagements
Nvidia
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Nvidia
Principal Software Engineer, Rack-Scale System Software - CSP Engagements
Nvidia
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole - fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements - including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software - ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software - health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior - champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices - drive documentation, tooling, and test strategy improvements as a result Collaborate with execution teams on left-shift strategy - ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams What we need to see: 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience) Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery) Understanding of error handling and recovery design patterns in distributed systems - fault isolation, retry policies, graceful degradation Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus Customer obsession - genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication - ability to translate complex system software architecture into actionable mentorship for customer engineering teams Ways to stand out from the crowd: Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms) Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices) Familiarity with GPU or accelerator fleet operations - driver lifecycle, firmware rollout strategies, health-based scheduling Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale NVIDIA's invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement. NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Kernel Engineer (Compute / Accelerator)
DensityAI Mountain View, California
About the role You will write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. This is the critical interface between high-level ML workloads and silicon - your code directly determines how effectively the hardware performs. You'll work closely with the architecture and compiler teams to define the kernel programming model, implement core tensor operations, and drive the performance profiling workflow that validates silicon design decisions. What you'll do Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation Develop and maintain profiling infrastructure to measure kernel performance against architectural targets Define and document shuffle patterns for ML kernel primitives across CPU-like control, tensor cores, and CUTLASS-style operations Drive kernel DSL design decisions - thread spawn mechanisms, register passing conventions, and memory management strategies Enable end-to-end kernel execution on the architectural simulator Collaborate with the compiler team on the MLIR dialect - your kernels are the primary validation target Create onboarding documentation and kernel writing guides for the broader team What we're looking for C/C++ - production-grade systems code, not scripted glue. You'll write performance-critical kernels CUDA or equivalent accelerator programming - deep experience writing GPU kernels, understanding warp/wavefront execution, memory coalescing, shared memory optimization. The mental model transfers directly Computer architecture - you need to reason about pipelines, memory hierarchies, data movement costs, and how software maps to hardware Performance profiling and optimization - you live in profilers. Identifying bottlenecks, measuring throughput, and iterating until kernels meet targets is the core loop Tensor operations - practical understanding of GEMM, convolution, attention, reduction, and scatter/gather as they map to hardware Python - for scripting, DSL integration, and profiling automation (Optional) RISC-V, x86, or ARM64 ISA experience (Optional) MLIR or LLVM compiler infrastructure (Optional) HPC or scientific computing background (large-scale parallel compute intuition) (Optional) FPGA or Verilog/SystemVerilog (ability to read RTL and reason about the hardware you're targeting) (Optional) Familiarity with CUTLASS, Triton, or similar kernel libraries Compensation Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO. Visa Sponsorship DensityAI sponsors qualified candidates for H-1B, O-1, TN, E-3, and other employment-based visas, and we welcome applicants on F-1 OPT and STEM-OPT. Work authorization is required at start; we provide immigration support to secure or transfer status. Equal Opportunity DensityAI is an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religious creed, national origin, ancestry, physical or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age (40+), sexual orientation, military or veteran status, pregnancy, or any other status protected by law. We comply with the California CROWN Act and provide reasonable accommodations on request. Full compensation packages are based on candidate experience and relevant certifications. California pay range $250,000-$320,000 USD Compensation Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO. Visa Sponsorship DensityAI sponsors qualified candidates for H-1B, O-1, TN, E-3, and other employment-based visas, and we welcome applicants on F-1 OPT and STEM-OPT. Work authorization is required at start; we provide immigration support to secure or transfer status. Export Controls Aspects of this role may involve access to information subject to U.S. export controls (EAR/ITAR). We may discuss licensing or scope adjustments during the interview. Equal Opportunity DensityAI is an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religious creed, national origin, ancestry, physical or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age (40+), sexual orientation, military or veteran status, pregnancy, or any other status protected by law. We comply with the California CROWN Act and provide reasonable accommodations on request.
09/25/2026
Full time
About the role You will write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. This is the critical interface between high-level ML workloads and silicon - your code directly determines how effectively the hardware performs. You'll work closely with the architecture and compiler teams to define the kernel programming model, implement core tensor operations, and drive the performance profiling workflow that validates silicon design decisions. What you'll do Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation Develop and maintain profiling infrastructure to measure kernel performance against architectural targets Define and document shuffle patterns for ML kernel primitives across CPU-like control, tensor cores, and CUTLASS-style operations Drive kernel DSL design decisions - thread spawn mechanisms, register passing conventions, and memory management strategies Enable end-to-end kernel execution on the architectural simulator Collaborate with the compiler team on the MLIR dialect - your kernels are the primary validation target Create onboarding documentation and kernel writing guides for the broader team What we're looking for C/C++ - production-grade systems code, not scripted glue. You'll write performance-critical kernels CUDA or equivalent accelerator programming - deep experience writing GPU kernels, understanding warp/wavefront execution, memory coalescing, shared memory optimization. The mental model transfers directly Computer architecture - you need to reason about pipelines, memory hierarchies, data movement costs, and how software maps to hardware Performance profiling and optimization - you live in profilers. Identifying bottlenecks, measuring throughput, and iterating until kernels meet targets is the core loop Tensor operations - practical understanding of GEMM, convolution, attention, reduction, and scatter/gather as they map to hardware Python - for scripting, DSL integration, and profiling automation (Optional) RISC-V, x86, or ARM64 ISA experience (Optional) MLIR or LLVM compiler infrastructure (Optional) HPC or scientific computing background (large-scale parallel compute intuition) (Optional) FPGA or Verilog/SystemVerilog (ability to read RTL and reason about the hardware you're targeting) (Optional) Familiarity with CUTLASS, Triton, or similar kernel libraries Compensation Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO. Visa Sponsorship DensityAI sponsors qualified candidates for H-1B, O-1, TN, E-3, and other employment-based visas, and we welcome applicants on F-1 OPT and STEM-OPT. Work authorization is required at start; we provide immigration support to secure or transfer status. Equal Opportunity DensityAI is an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religious creed, national origin, ancestry, physical or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age (40+), sexual orientation, military or veteran status, pregnancy, or any other status protected by law. We comply with the California CROWN Act and provide reasonable accommodations on request. Full compensation packages are based on candidate experience and relevant certifications. California pay range $250,000-$320,000 USD Compensation Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO. Visa Sponsorship DensityAI sponsors qualified candidates for H-1B, O-1, TN, E-3, and other employment-based visas, and we welcome applicants on F-1 OPT and STEM-OPT. Work authorization is required at start; we provide immigration support to secure or transfer status. Export Controls Aspects of this role may involve access to information subject to U.S. export controls (EAR/ITAR). We may discuss licensing or scope adjustments during the interview. Equal Opportunity DensityAI is an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religious creed, national origin, ancestry, physical or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age (40+), sexual orientation, military or veteran status, pregnancy, or any other status protected by law. We comply with the California CROWN Act and provide reasonable accommodations on request.
Nvidia
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence - CSP Engagements
Nvidia Santa Clara, California
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. In this role, you will augment NVIDIA's internal software/firmware and quality teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of reliability software/firmware architecture, methodology, incorporate their fleet telemetry and failure data into NVIDIA's improvement priorities, and validate that reliability improvements measured in the lab translate to real customer environments. Your cross-CSP visibility enables you to distinguish systemic architectural gaps from environmental or configuration-specific issues that no single customer engagement could identify alone. What you'll be doing: Drive reliability work streams with CSP engineering teams - ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architecture Gather and synthesize CSP fleet reliability data - identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teams Define consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practices Conduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specific Drive fleet health monitoring integration architecture - ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automation Define burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningful Collaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launch Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments What we need to see: 15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges. BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience) Deep expertise in multi-NUMA, rack-scale system software and firmware. Statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classification Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation Understanding of hardware failure modes in large-scale GPU/accelerator deployments - ability to classify and prioritize across compute, interconnect, memory, power, and thermal domains Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems. Familiarity with predictive maintenance or anomaly detection approaches applied to fleet health data Customer obsession - genuine passion for understanding fleet reliability challenges at scale and translating them into actionable engineering priorities Strong communication - ability to present statistical reliability findings to both deep technical audiences and executive leadership. Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority Ways to stand out from the crowd: Experience in fleet reliability at a hyperscaler (hardware health, fleet reliability at leading CSP/Hyperscaler) Familiarity with NVIDIA GPU error taxonomy (Xid errors, NVLink error counters, thermal events, CPER records) Experience building health scoring or predictive failure models for accelerator or HPC infrastructure Background in defining MTBI/MTBF measurement standards or certification programs for complex multi-component systems Understanding of how reliability data flows from device firmware through telemetry pipelines to fleet-level dashboards and automated remediation NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. In this role, you will augment NVIDIA's internal software/firmware and quality teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of reliability software/firmware architecture, methodology, incorporate their fleet telemetry and failure data into NVIDIA's improvement priorities, and validate that reliability improvements measured in the lab translate to real customer environments. Your cross-CSP visibility enables you to distinguish systemic architectural gaps from environmental or configuration-specific issues that no single customer engagement could identify alone. What you'll be doing: Drive reliability work streams with CSP engineering teams - ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architecture Gather and synthesize CSP fleet reliability data - identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teams Define consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practices Conduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specific Drive fleet health monitoring integration architecture - ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automation Define burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningful Collaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launch Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments What we need to see: 15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges. BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience) Deep expertise in multi-NUMA, rack-scale system software and firmware. Statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classification Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation Understanding of hardware failure modes in large-scale GPU/accelerator deployments - ability to classify and prioritize across compute, interconnect, memory, power, and thermal domains Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems. Familiarity with predictive maintenance or anomaly detection approaches applied to fleet health data Customer obsession - genuine passion for understanding fleet reliability challenges at scale and translating them into actionable engineering priorities Strong communication - ability to present statistical reliability findings to both deep technical audiences and executive leadership. Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority Ways to stand out from the crowd: Experience in fleet reliability at a hyperscaler (hardware health, fleet reliability at leading CSP/Hyperscaler) Familiarity with NVIDIA GPU error taxonomy (Xid errors, NVLink error counters, thermal events, CPER records) Experience building health scoring or predictive failure models for accelerator or HPC infrastructure Background in defining MTBI/MTBF measurement standards or certification programs for complex multi-component systems Understanding of how reliability data flows from device firmware through telemetry pipelines to fleet-level dashboards and automated remediation NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Senior Backend Engineer, Inference Platform
Together AI Remote, Oregon
About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the latest LLMs, multimodal models, image, audio, video, and speech models at scale. If you get a thrill from optimizing latency down to the last millisecond, this is your playground. You'll work hands-on with tens of thousands of GPUs (H100s, H200s, GB200s, and beyond), figuring out how to fully utilize every FLOP and every gigabyte of memory. You'll collaborate directly with research teams to bring frontier models into production, making breakthroughs usable in the real world. Our team also works closely with the open source community, contributing to and leveraging projects like SGLang, vLLM, and NVIDIA Dynamo to push the boundaries of inference performance and efficiency. Shape the core inference backbone that powers Together AI's frontier models. Solve performance-critical challenges in global request routing, load balancing, and large-scale resource allocation. Work with state-of-the-art accelerators (H100s, H200s, GB200s) at global scale. Partner with world-class researchers to bring new model architectures into production. Collaborate with and contribute to the open source community, shaping the tools that advance the industry. A culture of deep technical ownership and high impact - where your work makes models faster, cheaper, and more accessible. Competitive compensation, equity, and benefits. Responsibilities Build and optimize global and local request routing, ensuring low-latency load balancing across data centers and model engine pods. Develop auto-scaling systems to dynamically allocate resources and meet strict SLOs across dozens of data centers. Design systems for multi-tenant traffic shaping, tuning both resource allocation and request handling - including smart rate limiting and regulation - to ensure fairness and consistent experience across all users. Engineer trade-offs between latency and throughput to serve diverse workloads efficiently. Optimize prefix caching to reduce model compute and speed up responses. Collaborate with ML researchers to bring new model architectures into production at scale. Continuously profile and analyze system-level performance to identify bottlenecks and implement optimizations. Requirements 5+ years of demonstrated experience building large-scale, fault-tolerant, distributed systems and API microservices. Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems. Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance. Expert-level programming in one or more of: Rust, Go, Python, or TypeScript. Knowledge of modern LLMs and generative models and how they are served in production is a plus. Experience working with the open source ecosystem around inference is highly valuable; familiarity with SGLang, vLLM, or NVIDIA Dynamo will be especially handy. Experience with Kubernetes or container orchestration is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus. Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200,000 - $290,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
09/24/2026
Full time
About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the latest LLMs, multimodal models, image, audio, video, and speech models at scale. If you get a thrill from optimizing latency down to the last millisecond, this is your playground. You'll work hands-on with tens of thousands of GPUs (H100s, H200s, GB200s, and beyond), figuring out how to fully utilize every FLOP and every gigabyte of memory. You'll collaborate directly with research teams to bring frontier models into production, making breakthroughs usable in the real world. Our team also works closely with the open source community, contributing to and leveraging projects like SGLang, vLLM, and NVIDIA Dynamo to push the boundaries of inference performance and efficiency. Shape the core inference backbone that powers Together AI's frontier models. Solve performance-critical challenges in global request routing, load balancing, and large-scale resource allocation. Work with state-of-the-art accelerators (H100s, H200s, GB200s) at global scale. Partner with world-class researchers to bring new model architectures into production. Collaborate with and contribute to the open source community, shaping the tools that advance the industry. A culture of deep technical ownership and high impact - where your work makes models faster, cheaper, and more accessible. Competitive compensation, equity, and benefits. Responsibilities Build and optimize global and local request routing, ensuring low-latency load balancing across data centers and model engine pods. Develop auto-scaling systems to dynamically allocate resources and meet strict SLOs across dozens of data centers. Design systems for multi-tenant traffic shaping, tuning both resource allocation and request handling - including smart rate limiting and regulation - to ensure fairness and consistent experience across all users. Engineer trade-offs between latency and throughput to serve diverse workloads efficiently. Optimize prefix caching to reduce model compute and speed up responses. Collaborate with ML researchers to bring new model architectures into production at scale. Continuously profile and analyze system-level performance to identify bottlenecks and implement optimizations. Requirements 5+ years of demonstrated experience building large-scale, fault-tolerant, distributed systems and API microservices. Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems. Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance. Expert-level programming in one or more of: Rust, Go, Python, or TypeScript. Knowledge of modern LLMs and generative models and how they are served in production is a plus. Experience working with the open source ecosystem around inference is highly valuable; familiarity with SGLang, vLLM, or NVIDIA Dynamo will be especially handy. Experience with Kubernetes or container orchestration is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus. Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200,000 - $290,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
AI/HPC System Engineer
SK hynix America San Jose, California
About SK hynix America: At SK hynix America, we are at the forefront of semiconductor innovation. As a global leader in HBM, DRAM, and NAND flash technologies, we develop the advanced memory solutions powering everything from advanced mobile technology to massive AI data centers. With major investments in the U.S. and a leading position in the global AI revolution, we build the critical infrastructure of the digital landscape and remain committed to sustainable operations. We invite innovative minds to be part of our journey. Here, you will be part of a collaborative team pioneering the next generation of memory technologies, expanding our market footprint, and defining the future of computing. Work Model: Onsite Job Title: AI/HPC System Engineer Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite Job Summary: This role will be responsible for assuring SK hynix emerging technology leadership to be maintained and leading the next generation of memory-centric architecture using your expertise in computer system and SW technologies including open source tools to demonstrate benefits and use cases. AI/HPC System engineer will be responsible for designing, developing, and optimizing next-generation memory solutions and system software for energy-efficient AI computing platform. AI/HPC System engineer will work closely with our hardware and software teams to ensure seamless integration and optimal performance of our next-generation AI memory solutions Responsibilities: Individual technical contributor within the research team focused on next generation storage/memory systems. Design and develop next-generation memory solution prototype and system software for energy-efficient AI/ML computing infrastructure, including memory management, cache hierarchy, and data transfer protocols. Collaborate with hardware architects to ensure optimal HW/SW co-design and integration. Develop and maintain system software frameworks and APIs for AI system infrastructure. Optimize system software for performance, power efficiency, and scalability. Develop and maintain documentation, technical specifications, and test plans for system software. Participate in technical discussions and provide input on system software design and architecture with industry ecosystem partners Stay up-to-date with industry trends and emerging technologies in AI, memory systems, and system software. Verbal/written communication with both internal and external collaborator Qualifications: Ph.D. in Computer Science/Electrical Engineering or 6+ years of experience in system software development, with a focus on memory system architecture and AI infrastructure. Strong understanding of computer architecture, memory systems, and system software design. Experience and familiarity with Python/C/C++ programming languages Knowledge of operating systems, device drivers, and system programming. Good working knowledge or experience in GPU architecture and CUDA programming Excellent problem-solving skills, with the ability to debug complex system software issues. Strong communication and collaboration skills, with the ability to work with cross-functional teams. Benefits: Top Tier health insurance at no employee cost Paid day offs: Paid Time Off, Company Holidays, Parental Leave, Happy Fridays 401k Matching Flexible Spending Account (FSA) for Health Care & Dependent Care Educational reimbursement up to $10,000 per year Donation Matching and volunteering opportunities Corporate discount programs Free Breakfast/Lunch/Dinner provided to employees Compensation: Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. Pay within the provided range varies by work location and may also depend on job-related skills and experience. Your Recruiter can share more about the specific salary range for the job location during the hiring process. Pay Range $175,000-$230,000 USD Equal Employment Opportunity: SK hynix America is an Equal Employment Opportunity Employer. We provide equal employment opportunities to all qualified applicants and employees and prohibit discrimination and harassment of any type without regard to race, sex, pregnancy, sexual orientation, religion, age, gender identity, national origin, color, protected veteran or disability status, genetic information, or any other status protected under federal, state, or local applicable laws. In compliance with federal law, SK hynix America also participates in the federal E-Verify program. Upon accepting an offer of employment, all newly hired individuals will be required to complete Form I-9, which includes verifying identity and legal authorization to work in the United States. Please review the official government posters detailing your rights and protections prior to applying: E-Verify Participation Poster (Bilingual English & Spanish) IER Right to Work Poster (English) IER Right to Work Poster (Spanish)
09/24/2026
Full time
About SK hynix America: At SK hynix America, we are at the forefront of semiconductor innovation. As a global leader in HBM, DRAM, and NAND flash technologies, we develop the advanced memory solutions powering everything from advanced mobile technology to massive AI data centers. With major investments in the U.S. and a leading position in the global AI revolution, we build the critical infrastructure of the digital landscape and remain committed to sustainable operations. We invite innovative minds to be part of our journey. Here, you will be part of a collaborative team pioneering the next generation of memory technologies, expanding our market footprint, and defining the future of computing. Work Model: Onsite Job Title: AI/HPC System Engineer Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite Job Summary: This role will be responsible for assuring SK hynix emerging technology leadership to be maintained and leading the next generation of memory-centric architecture using your expertise in computer system and SW technologies including open source tools to demonstrate benefits and use cases. AI/HPC System engineer will be responsible for designing, developing, and optimizing next-generation memory solutions and system software for energy-efficient AI computing platform. AI/HPC System engineer will work closely with our hardware and software teams to ensure seamless integration and optimal performance of our next-generation AI memory solutions Responsibilities: Individual technical contributor within the research team focused on next generation storage/memory systems. Design and develop next-generation memory solution prototype and system software for energy-efficient AI/ML computing infrastructure, including memory management, cache hierarchy, and data transfer protocols. Collaborate with hardware architects to ensure optimal HW/SW co-design and integration. Develop and maintain system software frameworks and APIs for AI system infrastructure. Optimize system software for performance, power efficiency, and scalability. Develop and maintain documentation, technical specifications, and test plans for system software. Participate in technical discussions and provide input on system software design and architecture with industry ecosystem partners Stay up-to-date with industry trends and emerging technologies in AI, memory systems, and system software. Verbal/written communication with both internal and external collaborator Qualifications: Ph.D. in Computer Science/Electrical Engineering or 6+ years of experience in system software development, with a focus on memory system architecture and AI infrastructure. Strong understanding of computer architecture, memory systems, and system software design. Experience and familiarity with Python/C/C++ programming languages Knowledge of operating systems, device drivers, and system programming. Good working knowledge or experience in GPU architecture and CUDA programming Excellent problem-solving skills, with the ability to debug complex system software issues. Strong communication and collaboration skills, with the ability to work with cross-functional teams. Benefits: Top Tier health insurance at no employee cost Paid day offs: Paid Time Off, Company Holidays, Parental Leave, Happy Fridays 401k Matching Flexible Spending Account (FSA) for Health Care & Dependent Care Educational reimbursement up to $10,000 per year Donation Matching and volunteering opportunities Corporate discount programs Free Breakfast/Lunch/Dinner provided to employees Compensation: Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. Pay within the provided range varies by work location and may also depend on job-related skills and experience. Your Recruiter can share more about the specific salary range for the job location during the hiring process. Pay Range $175,000-$230,000 USD Equal Employment Opportunity: SK hynix America is an Equal Employment Opportunity Employer. We provide equal employment opportunities to all qualified applicants and employees and prohibit discrimination and harassment of any type without regard to race, sex, pregnancy, sexual orientation, religion, age, gender identity, national origin, color, protected veteran or disability status, genetic information, or any other status protected under federal, state, or local applicable laws. In compliance with federal law, SK hynix America also participates in the federal E-Verify program. Upon accepting an offer of employment, all newly hired individuals will be required to complete Form I-9, which includes verifying identity and legal authorization to work in the United States. Please review the official government posters detailing your rights and protections prior to applying: E-Verify Participation Poster (Bilingual English & Spanish) IER Right to Work Poster (English) IER Right to Work Poster (Spanish)

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board