What you'll do Own the strategy, requirements, priorities, and roadmap for a compute product area. Work directly with customers to identify workload needs and adoption barriers. Decide which opportunities warrant investment and which do not. Work with silicon partners and OCI engineering to evaluate processor and platform choices across performance, memory, interconnect, power, software readiness, cost, and delivery timing. Write decision papers for leadership that establish the context, assess alternatives, and recommend product investments, roadmap choices, and pricing. Make assumptions, economics, risks, and requested decisions explicit. Present and defend recommendations in leadership reviews. Resolve open questions and revise proposals when the evidence changes. Lead product decisions through development and launch, including scope tradeoffs, positioning, pricing, and readiness for customer use. Track adoption, workload competitiveness, and business performance. Use the results to change priorities and guide subsequent investments. Required experience and expertise Typically 8+ years of relevant industry experience, including 3+ years of product management experience with ownership of requirements, roadmap priorities, business tradeoffs, and product outcomes. Direct experience developing, evaluating, deploying, or bringing to market CPUs, GPUs, or AI accelerators, with technical depth in at least one category. Enough understanding of processor and system architecture to evaluate engineering proposals, challenge benchmark claims, and defend technical recommendations with experienced engineers. Prior product management responsibility for a processor, compute platform, or related infrastructure product. Your experience includes defining requirements, setting priorities, making business tradeoffs, and following a product through launch and adoption. Evidence of consequential product decisions you made: the alternatives considered, the supporting technical and customer evidence, and the results. Experience authoring decision papers or comparable proposals and defending recommendations with technical and business leaders. Ability to connect workload performance and system costs to pricing, customer value, and an investment case. Preferred experience Experience with cloud compute products or data center platforms. Experience working with processor vendors, system suppliers, or infrastructure customers on product requirements and roadmaps. Education A degree in electrical engineering, computer engineering, or computer science is preferred. Equivalent technical expertise demonstrated through relevant industry experience will also be considered. Years of experience are a guide. The scope of your ownership, quality of your judgment, and demonstrated results determine fit. Responsibilities Key Responsibilities Product Analysis - Market Analysis: -Prioritizes products/features based on evaluations of business value, feasibility, and user needs (internal and/or external). -Ensures alignment between the customer, the market, and Oracle's goals and strategy. -Analyzes market trends and the competitive landscape to maintain a competitive position. -Engages customers, non-customers, partners, and industry analysts to gather market information and identify opportunities for products. Product Analysis - Solution Identification: -Leads efforts to identify customers' unmet or unknown needs, market opportunities, and regulatory requirements. -Guides internal teams to define problem statements for complex products and develops hypotheses for new products. -Independently owns and drives the solutioning process by collaborating with stakeholders. -Drives evaluation and validation of solutions to inform decisions on enhancing products/features. -Creates, maintains, and reviews solution artifacts (e.g., regulatory, functional, and non-functional requirements). Product Ownership - Product Development and Roadmapping: -Takes ownership of planning, execution, and release processes. -Owns roadmaps for multiple products/features, ensuring alignment with business objectives. -Develops feature requirements, resources, and best practices for stakeholders. -Drives delivery excellence and resolves risks and issues. -Tracks and analyzes metrics (e.g., service requests, bugs, feedback) to make data-driven product decisions. -Writes and reviews user stories and acceptance criteria. -Determines feature and experience priorities to support OKRs (Objectives and Key Results). Product Ownership - Product Leadership and Influence: -Acts as an innovation leader, shaping product strategy and direction. -Guides product development to improve customer experience and achieve business goals. -Drives new feature development using agile and continuous-delivery methodologies. Alignment and Collaboration: -Shares insights with partners and ensures cross-team alignment. -Anticipates and manages stakeholder pushback using expertise and data. -Partners with engineering and senior teams to identify solutions linked to OKRs and KPIs. -Oversees business impact across organizations, managing success criteria and performance metrics. -Establishes parameters for user testing and product iteration. Customer-Centricity: -Serves as the primary customer advocate throughout the product lifecycle. -Actively engages with customers to gather feedback and insights. -Works with customers to align Oracle products to business goals. -Captures and analyzes customer requirements for new product capabilities and feature designs. -Ensures products are tailored to consumer needs during implementations. -Reviews and prioritizes customer issues for satisfaction and quality, coordinating resolutions. Product Release - Go-to-Market: -Develops release overviews, presentations, and demos for field teams. -Recruits and supports early adopters to validate solutions. -Contributes to the development of user documentation. -Defines and tracks product adoption metrics, identifying user behavior trends. -Presents demos to customers, partners, analysts, and field teams to promote value. Core Responsibilities Planning & Execution: -Manages and coordinates moderately complex tasks, monitoring timelines and deliverables for projects. -Delegates, monitors, and prioritizes work, providing technical oversight and adjusting plans as needed. Collaboration & Partnership: -Collaborates across the organization to align expectations and achieve objectives. -Leverages understanding of leaders, stakeholders, and customers to ensure solutions meet needs. -Supports inclusivity by seeking and respecting diverse perspectives. Problem Solving: -Identifies and addresses moderately complex issues by analyzing a range of data/information. -Proactively escalates unresolved or critical issues with assessments and suggested solutions. -Reviews, contributes to, and documents problem-solving strategies. Continuous Learning: -Pursues learning opportunities to expand knowledge, skills, and tools. -Stays updated with industry trends and best practices. -Seeks and uses ongoing feedback and training; coaches and mentors junior team members. Continuous Improvement: -Develops and recommends process improvements to increase efficiency and effectiveness. -Evaluates the impact on key stakeholders and solicits feedback for continued improvement. Performance and Development: -Contributes to the talent pipeline by participating in interviews, assessing candidates, and providing hiring recommendations. Qualifications Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $92,900 to $209,500 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9 . click apply for full job details
09/25/2026
Full time
What you'll do Own the strategy, requirements, priorities, and roadmap for a compute product area. Work directly with customers to identify workload needs and adoption barriers. Decide which opportunities warrant investment and which do not. Work with silicon partners and OCI engineering to evaluate processor and platform choices across performance, memory, interconnect, power, software readiness, cost, and delivery timing. Write decision papers for leadership that establish the context, assess alternatives, and recommend product investments, roadmap choices, and pricing. Make assumptions, economics, risks, and requested decisions explicit. Present and defend recommendations in leadership reviews. Resolve open questions and revise proposals when the evidence changes. Lead product decisions through development and launch, including scope tradeoffs, positioning, pricing, and readiness for customer use. Track adoption, workload competitiveness, and business performance. Use the results to change priorities and guide subsequent investments. Required experience and expertise Typically 8+ years of relevant industry experience, including 3+ years of product management experience with ownership of requirements, roadmap priorities, business tradeoffs, and product outcomes. Direct experience developing, evaluating, deploying, or bringing to market CPUs, GPUs, or AI accelerators, with technical depth in at least one category. Enough understanding of processor and system architecture to evaluate engineering proposals, challenge benchmark claims, and defend technical recommendations with experienced engineers. Prior product management responsibility for a processor, compute platform, or related infrastructure product. Your experience includes defining requirements, setting priorities, making business tradeoffs, and following a product through launch and adoption. Evidence of consequential product decisions you made: the alternatives considered, the supporting technical and customer evidence, and the results. Experience authoring decision papers or comparable proposals and defending recommendations with technical and business leaders. Ability to connect workload performance and system costs to pricing, customer value, and an investment case. Preferred experience Experience with cloud compute products or data center platforms. Experience working with processor vendors, system suppliers, or infrastructure customers on product requirements and roadmaps. Education A degree in electrical engineering, computer engineering, or computer science is preferred. Equivalent technical expertise demonstrated through relevant industry experience will also be considered. Years of experience are a guide. The scope of your ownership, quality of your judgment, and demonstrated results determine fit. Responsibilities Key Responsibilities Product Analysis - Market Analysis: -Prioritizes products/features based on evaluations of business value, feasibility, and user needs (internal and/or external). -Ensures alignment between the customer, the market, and Oracle's goals and strategy. -Analyzes market trends and the competitive landscape to maintain a competitive position. -Engages customers, non-customers, partners, and industry analysts to gather market information and identify opportunities for products. Product Analysis - Solution Identification: -Leads efforts to identify customers' unmet or unknown needs, market opportunities, and regulatory requirements. -Guides internal teams to define problem statements for complex products and develops hypotheses for new products. -Independently owns and drives the solutioning process by collaborating with stakeholders. -Drives evaluation and validation of solutions to inform decisions on enhancing products/features. -Creates, maintains, and reviews solution artifacts (e.g., regulatory, functional, and non-functional requirements). Product Ownership - Product Development and Roadmapping: -Takes ownership of planning, execution, and release processes. -Owns roadmaps for multiple products/features, ensuring alignment with business objectives. -Develops feature requirements, resources, and best practices for stakeholders. -Drives delivery excellence and resolves risks and issues. -Tracks and analyzes metrics (e.g., service requests, bugs, feedback) to make data-driven product decisions. -Writes and reviews user stories and acceptance criteria. -Determines feature and experience priorities to support OKRs (Objectives and Key Results). Product Ownership - Product Leadership and Influence: -Acts as an innovation leader, shaping product strategy and direction. -Guides product development to improve customer experience and achieve business goals. -Drives new feature development using agile and continuous-delivery methodologies. Alignment and Collaboration: -Shares insights with partners and ensures cross-team alignment. -Anticipates and manages stakeholder pushback using expertise and data. -Partners with engineering and senior teams to identify solutions linked to OKRs and KPIs. -Oversees business impact across organizations, managing success criteria and performance metrics. -Establishes parameters for user testing and product iteration. Customer-Centricity: -Serves as the primary customer advocate throughout the product lifecycle. -Actively engages with customers to gather feedback and insights. -Works with customers to align Oracle products to business goals. -Captures and analyzes customer requirements for new product capabilities and feature designs. -Ensures products are tailored to consumer needs during implementations. -Reviews and prioritizes customer issues for satisfaction and quality, coordinating resolutions. Product Release - Go-to-Market: -Develops release overviews, presentations, and demos for field teams. -Recruits and supports early adopters to validate solutions. -Contributes to the development of user documentation. -Defines and tracks product adoption metrics, identifying user behavior trends. -Presents demos to customers, partners, analysts, and field teams to promote value. Core Responsibilities Planning & Execution: -Manages and coordinates moderately complex tasks, monitoring timelines and deliverables for projects. -Delegates, monitors, and prioritizes work, providing technical oversight and adjusting plans as needed. Collaboration & Partnership: -Collaborates across the organization to align expectations and achieve objectives. -Leverages understanding of leaders, stakeholders, and customers to ensure solutions meet needs. -Supports inclusivity by seeking and respecting diverse perspectives. Problem Solving: -Identifies and addresses moderately complex issues by analyzing a range of data/information. -Proactively escalates unresolved or critical issues with assessments and suggested solutions. -Reviews, contributes to, and documents problem-solving strategies. Continuous Learning: -Pursues learning opportunities to expand knowledge, skills, and tools. -Stays updated with industry trends and best practices. -Seeks and uses ongoing feedback and training; coaches and mentors junior team members. Continuous Improvement: -Develops and recommends process improvements to increase efficiency and effectiveness. -Evaluates the impact on key stakeholders and solicits feedback for continued improvement. Performance and Development: -Contributes to the talent pipeline by participating in interviews, assessing candidates, and providing hiring recommendations. Qualifications Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $92,900 to $209,500 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9 . click apply for full job details
As a Principal Member of Technical Staff, you will own the software design and development for major components of Oracle's Cloud Infrastructure. You should be both a rock-solid lead developer, curious problem solver, a distributed systems generalist and/or skilled Linux engineer with Systems triage experiance able to dive deep into any part of the stack and low-level systems to design broad distributed system interactions. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn. This role resides within the Compute AI Infrastructure Bare Metal Provisioning team, which owns the critical infrastructure responsible for automating the full server lifecycle from new platform shape (AMD/Intel/Arm/Nvidia) creation, hardware bring-up to customer-ready instance provisioning and firmware management. The services operate at the intersection of bare metal hardware and full-stack orchestration frameworks, a unique combination where both distributed systems engineers and engineers with background in Linux and firmware are highly valued. The team interfaces directly with components like BMCs, NICs, SmartNICs, ILOMs, GPUs, and custom firmware stacks. The team builds high performance, scalable micro-services and tooling that provision, configure, secure, and validate server platforms across OCI's massive fleet of Compute and GPU Infrastructure. You will partner closely across other teams in Compute, Networking, Security, Data center Engineering, and Hardware Development to ensure OCI can launch, scale, and maintain new server platforms with minimal operational overhead and high reliability. You will work directly with cutting edge GPU hardware and see the direct impact of your work on the business. We strive for equity, inclusion, and respect for all. We are committed to the greater good in our products and our actions. We are constantly learning and taking opportunities to grow our careers and ourselves. We challenge each other to stretch beyond our past to build our future. You are the builder here. You will be part of a team of really smart, motivated, and diverse people and given the autonomy and support to do your best work. It is a dynamic and flexible workplace where you'll belong and be encouraged. If you are interested in building large-scale distributed infrastructure for the cloud, want to work on cutting edge GPU infrastructure and the latest Compute systems, have a knack for distributed systems and/or Linux development with Systems experiance then this is your team! Oracle is aggressively investing in the Oracle Cloud to provide the broadest, most comprehensive cloud in the industry. Responsibilities Job Responsibilities: You will own the software design and development for major components of Oracle's Cloud Infrastructure. You should be both a rock solid developer, driven problem solver and a distributed systems generalist and/or Linux developer with Systems experiance able to dive deep, design, develop, operate, and debug any part of the stack and low level systems such as Linux, Docker, Java web services and Terraform, as well as design broad distributed system interactions. You should have a tenacious attitude to improve the status quo, independently seek out problems to solve and take action to deliver results wherever needed. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn. Qualifications: 6-10+ years experience delivering and operating large scale, highly available distributed systems, Linux development and Systems debugging. Strong knowledge of Object Oriented programming such as C++ or Java, and experience with scripting languages such as Python. Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals. Experience with tools such as Terraform for Infrastructure as Code. Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures. Strong understanding of databases, NoSQL systems, storage and distributed persistence technologies. Strong troubleshooting and performance tuning skills. Experience building multi-tenant, virtualized infrastructure a strong plus. Qualifications Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9. 11 paid holidays 10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours. 11. Paid parental leave 12. Adoption assistance 13. Employee Stock Purchase Plan 14. Financial planning and group legal 15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted. As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing or by calling 1- in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
09/25/2026
Full time
As a Principal Member of Technical Staff, you will own the software design and development for major components of Oracle's Cloud Infrastructure. You should be both a rock-solid lead developer, curious problem solver, a distributed systems generalist and/or skilled Linux engineer with Systems triage experiance able to dive deep into any part of the stack and low-level systems to design broad distributed system interactions. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn. This role resides within the Compute AI Infrastructure Bare Metal Provisioning team, which owns the critical infrastructure responsible for automating the full server lifecycle from new platform shape (AMD/Intel/Arm/Nvidia) creation, hardware bring-up to customer-ready instance provisioning and firmware management. The services operate at the intersection of bare metal hardware and full-stack orchestration frameworks, a unique combination where both distributed systems engineers and engineers with background in Linux and firmware are highly valued. The team interfaces directly with components like BMCs, NICs, SmartNICs, ILOMs, GPUs, and custom firmware stacks. The team builds high performance, scalable micro-services and tooling that provision, configure, secure, and validate server platforms across OCI's massive fleet of Compute and GPU Infrastructure. You will partner closely across other teams in Compute, Networking, Security, Data center Engineering, and Hardware Development to ensure OCI can launch, scale, and maintain new server platforms with minimal operational overhead and high reliability. You will work directly with cutting edge GPU hardware and see the direct impact of your work on the business. We strive for equity, inclusion, and respect for all. We are committed to the greater good in our products and our actions. We are constantly learning and taking opportunities to grow our careers and ourselves. We challenge each other to stretch beyond our past to build our future. You are the builder here. You will be part of a team of really smart, motivated, and diverse people and given the autonomy and support to do your best work. It is a dynamic and flexible workplace where you'll belong and be encouraged. If you are interested in building large-scale distributed infrastructure for the cloud, want to work on cutting edge GPU infrastructure and the latest Compute systems, have a knack for distributed systems and/or Linux development with Systems experiance then this is your team! Oracle is aggressively investing in the Oracle Cloud to provide the broadest, most comprehensive cloud in the industry. Responsibilities Job Responsibilities: You will own the software design and development for major components of Oracle's Cloud Infrastructure. You should be both a rock solid developer, driven problem solver and a distributed systems generalist and/or Linux developer with Systems experiance able to dive deep, design, develop, operate, and debug any part of the stack and low level systems such as Linux, Docker, Java web services and Terraform, as well as design broad distributed system interactions. You should have a tenacious attitude to improve the status quo, independently seek out problems to solve and take action to deliver results wherever needed. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn. Qualifications: 6-10+ years experience delivering and operating large scale, highly available distributed systems, Linux development and Systems debugging. Strong knowledge of Object Oriented programming such as C++ or Java, and experience with scripting languages such as Python. Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals. Experience with tools such as Terraform for Infrastructure as Code. Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures. Strong understanding of databases, NoSQL systems, storage and distributed persistence technologies. Strong troubleshooting and performance tuning skills. Experience building multi-tenant, virtualized infrastructure a strong plus. Qualifications Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9. 11 paid holidays 10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours. 11. Paid parental leave 12. Adoption assistance 13. Employee Stock Purchase Plan 14. Financial planning and group legal 15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted. As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing or by calling 1- in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role: We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale Responsibilities: Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). Optimize latency and throughput of model inference under real production workloads. Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). Develop custom tools to trace, replay, and fix issues across the full stack - from orchestration down to GPU kernels. Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems. BASIC QUALIFICATIONS: Deep low-level systems programming (C/C++ or Rust) Experience with large-scale, high-concurrent production serving. Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). Strong background in system optimizations: batching, caching, load balancing, parallelism. Low-level inference optimizations: GPU kernels, code generation. Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics. Experience with testing, benchmarking, and reliability of inference services. Experience designing and implementing CI/CD infrastructure for inference. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
09/25/2026
Full time
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role: We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale Responsibilities: Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). Optimize latency and throughput of model inference under real production workloads. Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). Develop custom tools to trace, replay, and fix issues across the full stack - from orchestration down to GPU kernels. Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems. BASIC QUALIFICATIONS: Deep low-level systems programming (C/C++ or Rust) Experience with large-scale, high-concurrent production serving. Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). Strong background in system optimizations: batching, caching, load balancing, parallelism. Low-level inference optimizations: GPU kernels, code generation. Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics. Experience with testing, benchmarking, and reliability of inference services. Experience designing and implementing CI/CD infrastructure for inference. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
About the Role The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition to that, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems. As a Platform Engineer in Model Shaping, you will work at the intersection of backend engineering and infrastructure, building the foundational layers of Together's platform for model customization and evaluation. You will design, develop, and operate both the backend services and the underlying systems that enable us to sustainably and reliably scale production workflows launched by our users, as well as internal research experiments. You will operate in a cross-functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on. You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams. Responsibilities Design and build Together's systems and infrastructure for model customization, including user-facing features and internal improvements Contribute to reliability improvements for the platform, participating in an on-call rotation and improving processes for incident response Create and improve internal tooling for deployment, continuous integration, and observability Build a job orchestration platform spanning multiple datacenters, supporting a highly heterogeneous hardware landscape Partner with teams developing internal services, co-designing these services and incorporating them in systems built within Together Requirements 3+ years of experience in building infrastructure or backend components of production services Extensive experience designing, operating, and troubleshooting production Linux environments and Kubernetes-based platforms Strong software engineering background in Python or Go Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD) Cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare-metal/cloud environment Strong communication skills, be willing to document systems and processes and collaborate with peers of varying technical expertise Comfortable operating across the stack, from cluster operations and infrastructure automation to backend service development Experience in any of the following will make you stand out: Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters, ideally with hands-on experience in operating NVIDIA's networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA) Deployment of services for AI training or inference Networking fundamentals, including TCP/IP, DNS, routing, load balancing, TLS, and network debugging tools Maintaining or contributing to open-source projects About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
09/25/2026
Full time
About the Role The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition to that, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems. As a Platform Engineer in Model Shaping, you will work at the intersection of backend engineering and infrastructure, building the foundational layers of Together's platform for model customization and evaluation. You will design, develop, and operate both the backend services and the underlying systems that enable us to sustainably and reliably scale production workflows launched by our users, as well as internal research experiments. You will operate in a cross-functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on. You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams. Responsibilities Design and build Together's systems and infrastructure for model customization, including user-facing features and internal improvements Contribute to reliability improvements for the platform, participating in an on-call rotation and improving processes for incident response Create and improve internal tooling for deployment, continuous integration, and observability Build a job orchestration platform spanning multiple datacenters, supporting a highly heterogeneous hardware landscape Partner with teams developing internal services, co-designing these services and incorporating them in systems built within Together Requirements 3+ years of experience in building infrastructure or backend components of production services Extensive experience designing, operating, and troubleshooting production Linux environments and Kubernetes-based platforms Strong software engineering background in Python or Go Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD) Cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare-metal/cloud environment Strong communication skills, be willing to document systems and processes and collaborate with peers of varying technical expertise Comfortable operating across the stack, from cluster operations and infrastructure automation to backend service development Experience in any of the following will make you stand out: Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters, ideally with hands-on experience in operating NVIDIA's networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA) Deployment of services for AI training or inference Networking fundamentals, including TCP/IP, DNS, routing, load balancing, TLS, and network debugging tools Maintaining or contributing to open-source projects About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The AI Infrastructure team at Together AI is at the forefront of building and scaling the foundational systems that power our generative AI platform. The storage and observability team is crucial for designing, implementing, and maintaining robust distributed storage solutions, ensuring seamless data access and management. They are also responsible for developing comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively identifying and resolving issues. Responsibilities Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows. Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services. Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm. Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis. Define observability best practices. Requirements Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure). Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm). Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying. Deep understanding of containerization (Docker) and orchestration (Kubernetes). Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows. Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data. Preferred Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines. Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing. Contributions to open-source observability projects. Familiarity with security monitoring and compliance frameworks. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
09/25/2026
Full time
About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The AI Infrastructure team at Together AI is at the forefront of building and scaling the foundational systems that power our generative AI platform. The storage and observability team is crucial for designing, implementing, and maintaining robust distributed storage solutions, ensuring seamless data access and management. They are also responsible for developing comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively identifying and resolving issues. Responsibilities Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows. Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services. Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm. Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis. Define observability best practices. Requirements Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure). Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm). Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying. Deep understanding of containerization (Docker) and orchestration (Kubernetes). Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows. Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data. Preferred Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines. Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing. Contributions to open-source observability projects. Familiarity with security monitoring and compliance frameworks. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning. What you will be doing: Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters. Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components. Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware. What we need to see: M.S./Ph.D. degree in CS/CE or equivalent experience. 5+ years of relevant experience. Excellent C/C++ programming and debugging skills. Strong experience with Linux. Expert understanding of computer system architecture and operating systems. Experience with parallel programming interfaces and communication runtimes. Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment. Ways to stand out from the crowd: Deep understanding of technology and passionate about what you do. Experience with CUDA programming and NVIDIA GPUs. Knowledge of high-performance networks like InfiniBand, iWARP etc. Experience with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, etc. Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment. NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning. What you will be doing: Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters. Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components. Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware. What we need to see: M.S./Ph.D. degree in CS/CE or equivalent experience. 5+ years of relevant experience. Excellent C/C++ programming and debugging skills. Strong experience with Linux. Expert understanding of computer system architecture and operating systems. Experience with parallel programming interfaces and communication runtimes. Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment. Ways to stand out from the crowd: Deep understanding of technology and passionate about what you do. Experience with CUDA programming and NVIDIA GPUs. Knowledge of high-performance networks like InfiniBand, iWARP etc. Experience with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, etc. Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment. NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Influence the roadmap of communication libraries - NCCL & NVSHMEM. Collaborate with a very dynamic team across multiple time zones. What we need to see: B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 8+ software engineering and HPC/AI experience Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe) Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile) Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems) Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc) Adaptability and passion to learn new areas and tools Flexibility to work and communicate effectively across different teams and timezones Ways to stand out from the crowd: Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals) Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
(ID: ) Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH). Benefits We Offer: 100% Medical, Dental & Vision Coverage for Employees Paid Time Off and Paid Holidays 401K match up to 5% Educational Benefits for Career Growth Employee Referral Bonus Flexible Spending Accounts: Healthcare (FSA) Parking Reimbursement Account (PRK) Dependent Care Assistant Program (DCAP) Transportation Reimbursement Account (TRN) The Site Reliability Engineer role centers on modernizing and consolidating a complex multi-cloud environment across AWS, Azure, and GCP, building a scalable, secure, and observable platform from the ground up using Kubernetes, AI/ML infrastructure, and zero-trust principles. You'll combine DevOps and SRE practices to support mission-driven scientific and clinical programs, emphasizing automation, reliability, compliance, and proactive monitoring while enabling innovation through AI-driven tooling. The team culture is highly collaborative and growth-oriented, valuing experimentation, continuous learning, and cross-functional leadership, with opportunities to shape future multi-cloud and platform engineering solutions. Responsibilities: Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and Open-telemetry tools Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments Automate workload onboarding and offboarding processes, ensuring standardization and governance Track system ownership, dependencies, and lifecycle states for operational transparency Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP) Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards Develop "golden paths" and standardized platform templates for consistent workload deployment Automate provisioning, patching, configuration management, and environment lifecycle Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering Build developer-friendly platforms ("golden paths") that simplify deployments, reduce friction, and improve velocity Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads Build and manage containerized and orchestrated platforms (Docker, Kubernetes) Support cloud migration, modernization, and platform standardization initiatives Ensure systems meet security, compliance, backup, and disaster recovery requirements Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities Stay abreast of new technologies in your areas but not limited to AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering etc. Requirements: Must have total of 6+ experience DevOps / SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads. Must have 4+ years of Hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management and administration support Must have 4+ years of experience automating Infrastructure-as-Code (IaC) deployments to one of the following cloud platforms Amazon AWS, Google GCP and Microsoft Azure Must have 4+ years with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions Strong scripting skills (Python, Bash, PowerShell or similar) Must be proficient using vibe coding and coding assistants to develop scripts, tools and applications for the DevOps and SRE use cases Must have proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory Must be willing to learn new technologies, adopt and adapt to emerging technologies or needs from a project to a project Cloud certifications is preferred Certifications in Grafana, Splunk, Docker, Kubernetes is preferred but optional Disclaimer: The above description is meant to illustrate the general nature of work and level of effort being performed by individuals assigned to this position or job description. This is not restricted as a complete list of all skills, responsibilities, duties, and/or assignments required. Individuals may be required to perform duties outside of their position, job description or responsibilities as needed. The diversity of Axle's employees is a tremendous asset. We are firmly committed to providing equal opportunity in all aspects of employment and will not tolerate any illegal discrimination or harassment based on age, race, gender, religion, national origin, disability, marital status, covered veteran status, sexual orientation, status with respect to public assistance, and other characteristics protected under state, federal, or local law and to deter those who aid, abet, or induce discrimination or coerce others to discriminate. Accessibility: If you need an accommodation as part of the employment process please contact: This role has a market-competitive salary with an anticipated base compensation range listed below. Actual salaries will vary depending on a candidate's experience, qualifications, skills, and location. Salary Range $140,000-$155,000 USD
09/25/2026
Full time
(ID: ) Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH). Benefits We Offer: 100% Medical, Dental & Vision Coverage for Employees Paid Time Off and Paid Holidays 401K match up to 5% Educational Benefits for Career Growth Employee Referral Bonus Flexible Spending Accounts: Healthcare (FSA) Parking Reimbursement Account (PRK) Dependent Care Assistant Program (DCAP) Transportation Reimbursement Account (TRN) The Site Reliability Engineer role centers on modernizing and consolidating a complex multi-cloud environment across AWS, Azure, and GCP, building a scalable, secure, and observable platform from the ground up using Kubernetes, AI/ML infrastructure, and zero-trust principles. You'll combine DevOps and SRE practices to support mission-driven scientific and clinical programs, emphasizing automation, reliability, compliance, and proactive monitoring while enabling innovation through AI-driven tooling. The team culture is highly collaborative and growth-oriented, valuing experimentation, continuous learning, and cross-functional leadership, with opportunities to shape future multi-cloud and platform engineering solutions. Responsibilities: Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and Open-telemetry tools Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments Automate workload onboarding and offboarding processes, ensuring standardization and governance Track system ownership, dependencies, and lifecycle states for operational transparency Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP) Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards Develop "golden paths" and standardized platform templates for consistent workload deployment Automate provisioning, patching, configuration management, and environment lifecycle Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering Build developer-friendly platforms ("golden paths") that simplify deployments, reduce friction, and improve velocity Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads Build and manage containerized and orchestrated platforms (Docker, Kubernetes) Support cloud migration, modernization, and platform standardization initiatives Ensure systems meet security, compliance, backup, and disaster recovery requirements Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities Stay abreast of new technologies in your areas but not limited to AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering etc. Requirements: Must have total of 6+ experience DevOps / SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads. Must have 4+ years of Hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management and administration support Must have 4+ years of experience automating Infrastructure-as-Code (IaC) deployments to one of the following cloud platforms Amazon AWS, Google GCP and Microsoft Azure Must have 4+ years with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions Strong scripting skills (Python, Bash, PowerShell or similar) Must be proficient using vibe coding and coding assistants to develop scripts, tools and applications for the DevOps and SRE use cases Must have proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory Must be willing to learn new technologies, adopt and adapt to emerging technologies or needs from a project to a project Cloud certifications is preferred Certifications in Grafana, Splunk, Docker, Kubernetes is preferred but optional Disclaimer: The above description is meant to illustrate the general nature of work and level of effort being performed by individuals assigned to this position or job description. This is not restricted as a complete list of all skills, responsibilities, duties, and/or assignments required. Individuals may be required to perform duties outside of their position, job description or responsibilities as needed. The diversity of Axle's employees is a tremendous asset. We are firmly committed to providing equal opportunity in all aspects of employment and will not tolerate any illegal discrimination or harassment based on age, race, gender, religion, national origin, disability, marital status, covered veteran status, sexual orientation, status with respect to public assistance, and other characteristics protected under state, federal, or local law and to deter those who aid, abet, or induce discrimination or coerce others to discriminate. Accessibility: If you need an accommodation as part of the employment process please contact: This role has a market-competitive salary with an anticipated base compensation range listed below. Actual salaries will vary depending on a candidate's experience, qualifications, skills, and location. Salary Range $140,000-$155,000 USD
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
09/25/2026
Full time
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development, with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you'll be doing: Design and develop software solutions for data center servers including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-gen platforms. Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Collaborate with cross-functional teams in designing end-to-end solutions spanning firmware, OS, middleware, and applications with focus on AI/ML and HPC workloads. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Expert knowledge of Linux kernel internals, device drivers, communication protocols (PCIe, USB, Ethernet). Deep understanding of computer architecture, microprocessor concepts, and expert knowledge of ARM (aarch64) and x86 architectures. Ability to debug kernel crash dumps and all types of lock-up issues. Hands-on experience with GDB, kdump, and eBPF tracing to debug multiprocessor systems. Good understanding of ARM and Intel assembly. Strong understanding of PCIe virtualization and IOMMU. Deep understanding of NUMA architectures including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments. Strong programming skills in C/C++, Python, plus experience with virtualization, Kubernetes, and cloud-native architectures. Skilled in complex system-level debugging, performance analysis, and test design. BS or MS in Computer Engineering, Computer Science, or related field (or equivalent experience). 10+ years of system software development experience. Ways to stand out from the crowd: Experience with GPU computing (CUDA), deep learning workloads Expertise in Out of Band and In-band management architectures Knowledge of Memory fabric and CXL architectures NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you! Do you want to join a team of highly motivated and experienced program managers who drive the successful introduction of NVIDIA's next generation GPU/CPU based products? We work closely with internal leaders in Software, Hardware, Firmware, Marketing and Operations to ensure the SW team delivers outstanding products while operating across multiple functional units and all levels of management to achieve Time-To-Market. As part of the team, your knowledge of driver, firmware, diagnostics and the SW stack development processes and priorities will enable you to swiftly make the course adjustments needed to keep these complex projects on track! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
At Element Biosciences, we are passionate about our mission to empower the scientific community with more freedom and flexibility to accelerate our collective impact on humanity. We have built a highly efficient product-driven organization where employees can learn, grow, and thrive in a challenging but encouraging environment. We are committed to scientific integrity, collegiality, honesty, objectivity, and openness. We are seeking a highly skilled and motivated Senior Engineer, Machine Learning to join our dynamic team. The ideal candidate will have experience in data science and machine learning, with a background in working with multiomic and single-cell data and/or image processing and computer vision. This role involves creating, exploring, and analyzing models, as well as applying advanced image processing techniques to drive our research and development efforts and contribute to critical programs. This role will report to Vice President, AI and will be an onsite role at our headquarters in San Diego. If you possess the following and want to make a meaningful impact, we invite you to explore this role. Essential Functions and Responsibilities: Design, develop, and optimize deep learning models (CNNs, Vision Transformers, U-Net variants, and related architectures) for biological image analysis and classification Deploy and maintain production-grade neural network models on cloud infrastructure (e.g., AWS) or directly on imaging instruments, ensuring reliability, scalability, and performance Apply advanced image processing and computer vision techniques to analyze multimodal biological images, including segmentation, feature extraction, and quality scoring Develop and manage end-to-end ML pipelines - from data ingestion and preprocessing through model training, validation, and inference Analyze and interpret single-cell and multiomic data to support biological context and downstream interpretation of imaging results Collaborate with cross-functional teams including biology, software engineering, and instrumentation to co-design experiments and translate biological requirements into modeling objectives Explore and analyze large-scale imaging datasets to identify patterns, failure modes, and opportunities for model improvement Communicate findings, model performance metrics, and technical trade-offs to stakeholders through reports and presentations Stay current with the latest advances in deep learning, computer vision, and computational biology, and evaluate their applicability to internal research problems Education and Experience: Master's degree in Computer Science, Electrical Engineering, Bioinformatics, Computational Biology, or a related field with 5-7 years of relevant experience, or PhD with 0-3 years of experience Hands-on experience developing and deploying deep learning models for image analysis in production environments - either cloud-hosted or on-instrument - is required Strong proficiency with modern deep learning architectures including CNNs, Vision Transformers (ViT), U-Net, and attention-based models; familiarity with self-supervised or contrastive learning methods is a plus Experience with biological or biomedical image modalities (e.g., fluorescence microscopy, brightfield, high-content imaging) is strongly preferred Proficiency in Python and relevant deep learning and data science libraries: PyTorch, torchvision, OpenCV, Scikit-learn, NumPy, Pandas, and related tools Experience with cloud computing platforms (e.g., AWS), including model serving, containerization (Docker), and GPU-accelerated compute Familiarity with model calibration, uncertainty quantification, or performance evaluation frameworks is a plus Experience with single-cell or multiomic data analysis tools and workflows is a plus (not required) Knowledge of experimental design and statistical analysis Strong background in statistics and comfort reasoning about model outputs quantitatively Excellent problem-solving skills, attention to detail, and ability to work across scientific and engineering disciplines Physical Requirements: Frequently moves boxes weighing up to 20 pounds Location: San Diego - on-site Travel: Domestic travel up to 10% Job Type: Full-time/Exempt Base Compensation Pay Range: $139,000 - $183,000 In addition to base compensation noted above, you will be eligible for stock options, discretionary annual bonus, no cost health insurance plans, 401k with company match, and flexible paid time off. Please note: Base compensation will depend on multiple factors, including geographic location, qualifications, and experience. We foster an environment such that all people are afforded the freedom to pursue their passions without regard to race, color, religion, national or ethnic origin, gender (including pregnancy), sexual orientation, gender identity or expression, age, disability, veteran status or any other characteristics protected by law.
09/25/2026
Full time
At Element Biosciences, we are passionate about our mission to empower the scientific community with more freedom and flexibility to accelerate our collective impact on humanity. We have built a highly efficient product-driven organization where employees can learn, grow, and thrive in a challenging but encouraging environment. We are committed to scientific integrity, collegiality, honesty, objectivity, and openness. We are seeking a highly skilled and motivated Senior Engineer, Machine Learning to join our dynamic team. The ideal candidate will have experience in data science and machine learning, with a background in working with multiomic and single-cell data and/or image processing and computer vision. This role involves creating, exploring, and analyzing models, as well as applying advanced image processing techniques to drive our research and development efforts and contribute to critical programs. This role will report to Vice President, AI and will be an onsite role at our headquarters in San Diego. If you possess the following and want to make a meaningful impact, we invite you to explore this role. Essential Functions and Responsibilities: Design, develop, and optimize deep learning models (CNNs, Vision Transformers, U-Net variants, and related architectures) for biological image analysis and classification Deploy and maintain production-grade neural network models on cloud infrastructure (e.g., AWS) or directly on imaging instruments, ensuring reliability, scalability, and performance Apply advanced image processing and computer vision techniques to analyze multimodal biological images, including segmentation, feature extraction, and quality scoring Develop and manage end-to-end ML pipelines - from data ingestion and preprocessing through model training, validation, and inference Analyze and interpret single-cell and multiomic data to support biological context and downstream interpretation of imaging results Collaborate with cross-functional teams including biology, software engineering, and instrumentation to co-design experiments and translate biological requirements into modeling objectives Explore and analyze large-scale imaging datasets to identify patterns, failure modes, and opportunities for model improvement Communicate findings, model performance metrics, and technical trade-offs to stakeholders through reports and presentations Stay current with the latest advances in deep learning, computer vision, and computational biology, and evaluate their applicability to internal research problems Education and Experience: Master's degree in Computer Science, Electrical Engineering, Bioinformatics, Computational Biology, or a related field with 5-7 years of relevant experience, or PhD with 0-3 years of experience Hands-on experience developing and deploying deep learning models for image analysis in production environments - either cloud-hosted or on-instrument - is required Strong proficiency with modern deep learning architectures including CNNs, Vision Transformers (ViT), U-Net, and attention-based models; familiarity with self-supervised or contrastive learning methods is a plus Experience with biological or biomedical image modalities (e.g., fluorescence microscopy, brightfield, high-content imaging) is strongly preferred Proficiency in Python and relevant deep learning and data science libraries: PyTorch, torchvision, OpenCV, Scikit-learn, NumPy, Pandas, and related tools Experience with cloud computing platforms (e.g., AWS), including model serving, containerization (Docker), and GPU-accelerated compute Familiarity with model calibration, uncertainty quantification, or performance evaluation frameworks is a plus Experience with single-cell or multiomic data analysis tools and workflows is a plus (not required) Knowledge of experimental design and statistical analysis Strong background in statistics and comfort reasoning about model outputs quantitatively Excellent problem-solving skills, attention to detail, and ability to work across scientific and engineering disciplines Physical Requirements: Frequently moves boxes weighing up to 20 pounds Location: San Diego - on-site Travel: Domestic travel up to 10% Job Type: Full-time/Exempt Base Compensation Pay Range: $139,000 - $183,000 In addition to base compensation noted above, you will be eligible for stock options, discretionary annual bonus, no cost health insurance plans, 401k with company match, and flexible paid time off. Please note: Base compensation will depend on multiple factors, including geographic location, qualifications, and experience. We foster an environment such that all people are afforded the freedom to pursue their passions without regard to race, color, religion, national or ethnic origin, gender (including pregnancy), sexual orientation, gender identity or expression, age, disability, veteran status or any other characteristics protected by law.
About Marvell Marvell's semiconductor solutions are the essential building blocks of the data infrastructure that connects our world. Across enterprise, cloud and AI, and carrier architectures, our innovative technology is enabling new possibilities. At Marvell, you can affect the arc of individual lives, lift the trajectory of entire industries, and fuel the transformative potential of tomorrow. For those looking to make their mark on purposeful and enduring innovation, above and beyond fleeting trends, Marvell is a place to thrive, learn, and lead. Your Team, Your Impact The Technology and Solutions Architecture team is innovating at the intersection of hardware and software- Building co-optimized solutions for the most demanding AI data center challenges, reimagining the data- movement / storage / transformation / processing pipelines with purpose-built Marvell accelerators to advance next-generation AI infrastructure, and shaping open standards We are seeking an innovative AI Solutions Analyst to research, explore, and build co-optimized fabric-attached, shared, disaggregated, and multi-tier memory and storage solution architectures designed to meet the rapidly scaling memory and storage demands of XPU for increasingly data-centric workloads. You will work closely with architects, technologists, and product teams to influence research and development in cutting-edge forward-looking technologies in the disaggregated memory, storage, networking, and security domains. What You Can Expect Develop Next-Gen Solutions: Partner with hardware and software architects, technologists, and IC engineering teams to develop scalable, high-performance, cutting-edge AI/ML infrastructure. Foster Innovation: Spearhead advancements in accelerating AI/ML applications in heterogenous and disaggregated compute architectures, networking, storage, and security domains, focusing on emerging technologies such as CXL, UAL, UEC, low-latency transport protocols, and Post-Quantum Cryptography. Architect and Execute: Lead system architectural design and implement proof-of-concept solutions for diverse workloads, optimizing execution across distributed, heterogeneous memory, compute, and storage environments. Present innovative ideas and substantiated proof-points to senior management and CTOs. Represent in Standards Bodies: Advocate Marvell's leadership in industry standards bodies and working groups, shaping new standards and emerging technologies aligned with Marvell's strategic interests. Analyze Performance: Conduct advanced characterization and performance analysis of complex workloads, monitoring critical metrics such as response time, latency, and resource utilization. Build Simulation Environments: Create and refine simulation environments for large-scale setups and configurations, especially in the AI/ML domain. Document and Visualize Insights: Develop comprehensive documentation and data visualizations to support analysis and insights, leveraging tools like Python and related libraries. Stay at the Forefront: Keep abreast of the latest developments in generative AI and hardware data analysis, and propose innovative approaches to enhance device tuning and interoperability performance. What We're Looking For PhD or Master's degree in computer science and engineering, electrical engineering, or related field HW/SW co-design experience with proven expertise with GPUs, DPUs, FPGAs, or custom hardware accelerators Processor micro-architectures (preferably ARM), scalar and vector processing pipelines, techniques to optimize memory bound applications Deep understanding of memory or storage subsystems, disaggregated memory for accelerators, cluster level shared memory approaches, NVMeoF Familiarity with coherent / non-coherent interconnects, like CXL, UAL, NVLink, and multi-GPU / multi-node communication mechanisms and optimization techniques like SHARP and NCCL Proficiency in performance benchmarking of AI data processing workloads, system profiling and analyzing end-to-end flows Experience with simulators & modeling, GPU/CPU cooperative accelerated computing optimizing / accelerating data-centric AI workloads Expected Base Pay Range (USD) 94,160 - 141,000, $ per annum The successful candidate's starting base pay will be determined based on job-related skills, experience, qualifications, work location and market conditions. The expected base pay range for this role may be modified based on market conditions. This role is eligible to participate in Marvell's bonus and equity plans, including a new hire equity grant and annual equity refreshers, ensuring you are rewarded for your contributions and remain invested in the company's long-term success. Additional Compensation and Benefit Elements Marvell is committed to providing exceptional, comprehensive benefits that support our employees at every stage - from internship to retirement and through life's most important moments. Our offerings are built around four key pillars: financial well-being, family support, mental and physical health, and recognition. Highlights include an employee stock purchase plan with a 2-year look back, family support programs to help balance work and home life, robust mental health resources to prioritize emotional well-being, and a recognition and service awards to celebrate contributions and milestones. We look forward to sharing more with you during the interview process. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Any applicant who requires a reasonable accommodation during the selection process should contact Marvell HR Helpdesk at . Interview Integrity To support fair and authentic hiring practices, candidates are not permitted to use AI tools (such as transcription apps, real-time answer generators like ChatGPT or Copilot, or automated note-taking bots) during interviews. These tools must not be used to record, assist with, or enhance responses in any way. Our interviews are designed to evaluate your individual experience, thought process, and communication skills in real time. Use of AI tools without prior instruction from the interviewer will result in disqualification from the hiring process. This position may require access to technology and/or software subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). As such, applicants must be eligible to access export-controlled information as defined under applicable law. Marvell may be required to obtain export licensing approval from the U.S. Department of Commerce and/or the U.S. Department of State. Except for U.S. citizens, lawful permanent residents, or protected individuals as defined by 8 U.S.C. 1324b(a)(3), all applicants may be subject to an export license review process prior to employment.
09/25/2026
Full time
About Marvell Marvell's semiconductor solutions are the essential building blocks of the data infrastructure that connects our world. Across enterprise, cloud and AI, and carrier architectures, our innovative technology is enabling new possibilities. At Marvell, you can affect the arc of individual lives, lift the trajectory of entire industries, and fuel the transformative potential of tomorrow. For those looking to make their mark on purposeful and enduring innovation, above and beyond fleeting trends, Marvell is a place to thrive, learn, and lead. Your Team, Your Impact The Technology and Solutions Architecture team is innovating at the intersection of hardware and software- Building co-optimized solutions for the most demanding AI data center challenges, reimagining the data- movement / storage / transformation / processing pipelines with purpose-built Marvell accelerators to advance next-generation AI infrastructure, and shaping open standards We are seeking an innovative AI Solutions Analyst to research, explore, and build co-optimized fabric-attached, shared, disaggregated, and multi-tier memory and storage solution architectures designed to meet the rapidly scaling memory and storage demands of XPU for increasingly data-centric workloads. You will work closely with architects, technologists, and product teams to influence research and development in cutting-edge forward-looking technologies in the disaggregated memory, storage, networking, and security domains. What You Can Expect Develop Next-Gen Solutions: Partner with hardware and software architects, technologists, and IC engineering teams to develop scalable, high-performance, cutting-edge AI/ML infrastructure. Foster Innovation: Spearhead advancements in accelerating AI/ML applications in heterogenous and disaggregated compute architectures, networking, storage, and security domains, focusing on emerging technologies such as CXL, UAL, UEC, low-latency transport protocols, and Post-Quantum Cryptography. Architect and Execute: Lead system architectural design and implement proof-of-concept solutions for diverse workloads, optimizing execution across distributed, heterogeneous memory, compute, and storage environments. Present innovative ideas and substantiated proof-points to senior management and CTOs. Represent in Standards Bodies: Advocate Marvell's leadership in industry standards bodies and working groups, shaping new standards and emerging technologies aligned with Marvell's strategic interests. Analyze Performance: Conduct advanced characterization and performance analysis of complex workloads, monitoring critical metrics such as response time, latency, and resource utilization. Build Simulation Environments: Create and refine simulation environments for large-scale setups and configurations, especially in the AI/ML domain. Document and Visualize Insights: Develop comprehensive documentation and data visualizations to support analysis and insights, leveraging tools like Python and related libraries. Stay at the Forefront: Keep abreast of the latest developments in generative AI and hardware data analysis, and propose innovative approaches to enhance device tuning and interoperability performance. What We're Looking For PhD or Master's degree in computer science and engineering, electrical engineering, or related field HW/SW co-design experience with proven expertise with GPUs, DPUs, FPGAs, or custom hardware accelerators Processor micro-architectures (preferably ARM), scalar and vector processing pipelines, techniques to optimize memory bound applications Deep understanding of memory or storage subsystems, disaggregated memory for accelerators, cluster level shared memory approaches, NVMeoF Familiarity with coherent / non-coherent interconnects, like CXL, UAL, NVLink, and multi-GPU / multi-node communication mechanisms and optimization techniques like SHARP and NCCL Proficiency in performance benchmarking of AI data processing workloads, system profiling and analyzing end-to-end flows Experience with simulators & modeling, GPU/CPU cooperative accelerated computing optimizing / accelerating data-centric AI workloads Expected Base Pay Range (USD) 94,160 - 141,000, $ per annum The successful candidate's starting base pay will be determined based on job-related skills, experience, qualifications, work location and market conditions. The expected base pay range for this role may be modified based on market conditions. This role is eligible to participate in Marvell's bonus and equity plans, including a new hire equity grant and annual equity refreshers, ensuring you are rewarded for your contributions and remain invested in the company's long-term success. Additional Compensation and Benefit Elements Marvell is committed to providing exceptional, comprehensive benefits that support our employees at every stage - from internship to retirement and through life's most important moments. Our offerings are built around four key pillars: financial well-being, family support, mental and physical health, and recognition. Highlights include an employee stock purchase plan with a 2-year look back, family support programs to help balance work and home life, robust mental health resources to prioritize emotional well-being, and a recognition and service awards to celebrate contributions and milestones. We look forward to sharing more with you during the interview process. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Any applicant who requires a reasonable accommodation during the selection process should contact Marvell HR Helpdesk at . Interview Integrity To support fair and authentic hiring practices, candidates are not permitted to use AI tools (such as transcription apps, real-time answer generators like ChatGPT or Copilot, or automated note-taking bots) during interviews. These tools must not be used to record, assist with, or enhance responses in any way. Our interviews are designed to evaluate your individual experience, thought process, and communication skills in real time. Use of AI tools without prior instruction from the interviewer will result in disqualification from the hiring process. This position may require access to technology and/or software subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). As such, applicants must be eligible to access export-controlled information as defined under applicable law. Marvell may be required to obtain export licensing approval from the U.S. Department of Commerce and/or the U.S. Department of State. Except for U.S. citizens, lawful permanent residents, or protected individuals as defined by 8 U.S.C. 1324b(a)(3), all applicants may be subject to an export license review process prior to employment.
"In 36 months, agentic AI systems will be an operating reality across major institutions. We intend to be central to it." - Dr. Jim Rebesco, Cofounder and CEO, Striveworks The government's demand for AI is growing far faster than the systems required to support it. Fewer than 15% of federal AI programs have reached sustained production, despite billions of dollars invested. The models perform in testing, but they degrade in the real world. And when performance drops, trust goes with it. Striveworks was built to solve that problem. What you'll build Since 2018, we have delivered the most trusted AI systems operating in real-world use cases-providing a layer of assurance underneath hundreds of deployed models that monitors performance, manages drift, and sustains systems long after they leave the lab. As a Senior DevOps Engineer based at Schofield Barracks, you will own what you deploy. That means maintaining, optimizing, and enhancing the environments that national security customers depend on. You'll deploy custom K8s clusters, manage IaC automation, and serve as the operational link between platform developers and customer-facing teams. In this role, you will bring junior DevOps teammates along as you go, and you will have the latitude to explore better solutions when you find them. What it's like here We lead with trust, treat each other with respect, and use candor consistently, kindly, and constructively. We care deeply about our work, and we find genuine satisfaction in doing it well. Above all, we take ownership-because we feel the weight of collective results personally. We are looking for people who share these values and are eager to put them into practice. What we're looking for 6+ years of direct, hands-on experience in: Microservices deployment in Kubernetes Diagnosing and resolving issues within containerized environments Helm Chart and Kustomizations development/deployment Python and/or Golang programming, or other general purpose programming languages Automation and IaC, e.g., Terraform Cloud infrastructure: AWS and Azure Managing and troubleshooting Linux systems, e.g., networking and bash scripting Active Secret (or above) US security clearance and US citizenship The following isn't required, but we'd love to see it: Experience with DOD networking, tools, infrastructure, security requirements, and policies Proficiency with US federal information system security policies including STIGs, NIST SP 800-171, NIST SP 800-53, CMMC, and ICD 503 Experience with software deployments to unclassified, CUI, and classified DOD networks Experience with DevSecOps and CI/CD for GPU-enabled server administration and deployment Experience deploying or maintaining CNCF projects Experience with NAS and SAN technologies Experience with K8s and cloud-native applications in DDIL environments This role is hybrid/on site at customer sites at Schofield Barracks in Oahu, Hawaii. You will be expected to travel up to 30% of the time, including some international travel. Compensation The anticipated base pay range for this position is $220,000-$270,000/year. Striveworks' total compensation package includes a competitive base salary, equity grants, and cash bonuses. Benefits include: Medical/dental/vision insurance Voluntary life, long-term disability, accident, and hospital indemnity insurance HSA and FSA (including dependent care FSA) plans 401(k) plan Unlimited PTO Paid parental leave Ready to build systems that work for a mission that matters? Let's talk. Striveworks is an Equal Opportunity Employer and does not discriminate in employment on the basis of race, color, religion, belief, sex (including pregnancy and gender identity or expression), national origin, social or ethnic origin, political affiliation, sexual orientation, marital status, disability, genetic information, age, membership in an employee organization, retaliation, parental status, military service, or other non-merit factors. Striveworks will not tolerate discrimination or harassment of any kind. If you require assistance or a reasonable accommodation in the application process, please contact People Operations at . In compliance with federal law, all persons hired will be required to verify their identity and eligibility to work in the United States and to complete an employment eligibility verification form upon hire. Striveworks is a participating employer in the E-Verify program.
09/25/2026
Full time
"In 36 months, agentic AI systems will be an operating reality across major institutions. We intend to be central to it." - Dr. Jim Rebesco, Cofounder and CEO, Striveworks The government's demand for AI is growing far faster than the systems required to support it. Fewer than 15% of federal AI programs have reached sustained production, despite billions of dollars invested. The models perform in testing, but they degrade in the real world. And when performance drops, trust goes with it. Striveworks was built to solve that problem. What you'll build Since 2018, we have delivered the most trusted AI systems operating in real-world use cases-providing a layer of assurance underneath hundreds of deployed models that monitors performance, manages drift, and sustains systems long after they leave the lab. As a Senior DevOps Engineer based at Schofield Barracks, you will own what you deploy. That means maintaining, optimizing, and enhancing the environments that national security customers depend on. You'll deploy custom K8s clusters, manage IaC automation, and serve as the operational link between platform developers and customer-facing teams. In this role, you will bring junior DevOps teammates along as you go, and you will have the latitude to explore better solutions when you find them. What it's like here We lead with trust, treat each other with respect, and use candor consistently, kindly, and constructively. We care deeply about our work, and we find genuine satisfaction in doing it well. Above all, we take ownership-because we feel the weight of collective results personally. We are looking for people who share these values and are eager to put them into practice. What we're looking for 6+ years of direct, hands-on experience in: Microservices deployment in Kubernetes Diagnosing and resolving issues within containerized environments Helm Chart and Kustomizations development/deployment Python and/or Golang programming, or other general purpose programming languages Automation and IaC, e.g., Terraform Cloud infrastructure: AWS and Azure Managing and troubleshooting Linux systems, e.g., networking and bash scripting Active Secret (or above) US security clearance and US citizenship The following isn't required, but we'd love to see it: Experience with DOD networking, tools, infrastructure, security requirements, and policies Proficiency with US federal information system security policies including STIGs, NIST SP 800-171, NIST SP 800-53, CMMC, and ICD 503 Experience with software deployments to unclassified, CUI, and classified DOD networks Experience with DevSecOps and CI/CD for GPU-enabled server administration and deployment Experience deploying or maintaining CNCF projects Experience with NAS and SAN technologies Experience with K8s and cloud-native applications in DDIL environments This role is hybrid/on site at customer sites at Schofield Barracks in Oahu, Hawaii. You will be expected to travel up to 30% of the time, including some international travel. Compensation The anticipated base pay range for this position is $220,000-$270,000/year. Striveworks' total compensation package includes a competitive base salary, equity grants, and cash bonuses. Benefits include: Medical/dental/vision insurance Voluntary life, long-term disability, accident, and hospital indemnity insurance HSA and FSA (including dependent care FSA) plans 401(k) plan Unlimited PTO Paid parental leave Ready to build systems that work for a mission that matters? Let's talk. Striveworks is an Equal Opportunity Employer and does not discriminate in employment on the basis of race, color, religion, belief, sex (including pregnancy and gender identity or expression), national origin, social or ethnic origin, political affiliation, sexual orientation, marital status, disability, genetic information, age, membership in an employee organization, retaliation, parental status, military service, or other non-merit factors. Striveworks will not tolerate discrimination or harassment of any kind. If you require assistance or a reasonable accommodation in the application process, please contact People Operations at . In compliance with federal law, all persons hired will be required to verify their identity and eligibility to work in the United States and to complete an employment eligibility verification form upon hire. Striveworks is a participating employer in the E-Verify program.
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . What You'll Do: The Storage Engine Team at CoreWeave is responsible for the product capabilities and data plane function of CoreWeave's managed storage products. We build reliable, scalable storage solutions with segment leading performance. Storage engine works with engineering teams across infrastructure, compute, and platform to ensure our storage services meet the needs of the world's most demanding AI workloads. About the role: Design and Implement distributed storage solutions to support scaling data intensive AI workloads. Contribute to the development of exabyte-scale, S3-compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. Lead efforts to improve the reliability, durability, security, and observability of our storage stack. Collaborate with operations teams to monitor, troubleshoot, and improve storage systems in production environments. Set the bar for developing metrics and dashboards to provide visibility into storage performance and health. Analyze telemetry and system data to drive improvements in throughput, latency, and resilience. Work cross-functionally with platform, product, and infrastructure teams to deliver seamless storage capabilities across the stack. Share your knowledge and mentor other engineers on best practices in building distributed, high-performance systems. Who You Are: Bachelor's, Master's, or PhD degree in Computer Science, Engineering, or a related field. 8-10+ years of experience working in storage systems engineering or infrastructure. Strong hands-on experience with object storage or distributed filesystems in production environments. Experience with one or more storage protocols (e.g. S3, NFS) and file systems such as Ceph, DAOS, or similar. Proficiency in a systems programming language such as Go, C, or Rust. Proficiency leveraging AI tools to augment software development. Familiarity with storage observability tools and telemetry pipelines (e.g., ClickHouse, Prometheus, Grafana). Experience working with cloud-native infrastructure, Kubernetes, and scalable system architectures. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You love to grow and push the boundaries of your expertise You're curious about distributed storage and the demands AI/ML workloads You're an expert in data persistence on physical media, high performance data transfer using RDMA, or resilient distributed systems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $143,000 to $210,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) , CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: . Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C. 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.
09/25/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . What You'll Do: The Storage Engine Team at CoreWeave is responsible for the product capabilities and data plane function of CoreWeave's managed storage products. We build reliable, scalable storage solutions with segment leading performance. Storage engine works with engineering teams across infrastructure, compute, and platform to ensure our storage services meet the needs of the world's most demanding AI workloads. About the role: Design and Implement distributed storage solutions to support scaling data intensive AI workloads. Contribute to the development of exabyte-scale, S3-compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. Lead efforts to improve the reliability, durability, security, and observability of our storage stack. Collaborate with operations teams to monitor, troubleshoot, and improve storage systems in production environments. Set the bar for developing metrics and dashboards to provide visibility into storage performance and health. Analyze telemetry and system data to drive improvements in throughput, latency, and resilience. Work cross-functionally with platform, product, and infrastructure teams to deliver seamless storage capabilities across the stack. Share your knowledge and mentor other engineers on best practices in building distributed, high-performance systems. Who You Are: Bachelor's, Master's, or PhD degree in Computer Science, Engineering, or a related field. 8-10+ years of experience working in storage systems engineering or infrastructure. Strong hands-on experience with object storage or distributed filesystems in production environments. Experience with one or more storage protocols (e.g. S3, NFS) and file systems such as Ceph, DAOS, or similar. Proficiency in a systems programming language such as Go, C, or Rust. Proficiency leveraging AI tools to augment software development. Familiarity with storage observability tools and telemetry pipelines (e.g., ClickHouse, Prometheus, Grafana). Experience working with cloud-native infrastructure, Kubernetes, and scalable system architectures. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You love to grow and push the boundaries of your expertise You're curious about distributed storage and the demands AI/ML workloads You're an expert in data persistence on physical media, high performance data transfer using RDMA, or resilient distributed systems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $143,000 to $210,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) , CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: . Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C. 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . What You'll Do: The Storage Engine Team at CoreWeave is responsible for the product capabilities and data plane function of CoreWeave's managed storage products. We build reliable, scalable storage solutions with segment leading performance. Storage engine works with engineering teams across infrastructure, compute, and platform to ensure our storage services meet the needs of the world's most demanding AI workloads. About the role: Design and Implement distributed storage solutions to support scaling data intensive AI workloads. Contribute to the development of exabyte-scale, S3-compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. Lead efforts to improve the reliability, durability, security, and observability of our storage stack. Collaborate with operations teams to monitor, troubleshoot, and improve storage systems in production environments. Set the bar for developing metrics and dashboards to provide visibility into storage performance and health. Analyze telemetry and system data to drive improvements in throughput, latency, and resilience. Work cross-functionally with platform, product, and infrastructure teams to deliver seamless storage capabilities across the stack. Share your knowledge and mentor other engineers on best practices in building distributed, high-performance systems. Who You Are: Bachelor's, Master's, or PhD degree in Computer Science, Engineering, or a related field. 8-10+ years of experience working in storage systems engineering or infrastructure. Strong hands-on experience with object storage or distributed filesystems in production environments. Experience with one or more storage protocols (e.g. S3, NFS) and file systems such as Ceph, DAOS, or similar. Proficiency in a systems programming language such as Go, C, or Rust. Proficiency leveraging AI tools to augment software development. Familiarity with storage observability tools and telemetry pipelines (e.g., ClickHouse, Prometheus, Grafana). Experience working with cloud-native infrastructure, Kubernetes, and scalable system architectures. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You love to grow and push the boundaries of your expertise You're curious about distributed storage and the demands AI/ML workloads You're an expert in data persistence on physical media, high performance data transfer using RDMA, or resilient distributed systems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $143,000 to $210,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) , CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: . Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C. 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.
09/25/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . What You'll Do: The Storage Engine Team at CoreWeave is responsible for the product capabilities and data plane function of CoreWeave's managed storage products. We build reliable, scalable storage solutions with segment leading performance. Storage engine works with engineering teams across infrastructure, compute, and platform to ensure our storage services meet the needs of the world's most demanding AI workloads. About the role: Design and Implement distributed storage solutions to support scaling data intensive AI workloads. Contribute to the development of exabyte-scale, S3-compatible object storage and integrate dedicated storage clusters into diverse customer environments. Work with technologies such as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. Lead efforts to improve the reliability, durability, security, and observability of our storage stack. Collaborate with operations teams to monitor, troubleshoot, and improve storage systems in production environments. Set the bar for developing metrics and dashboards to provide visibility into storage performance and health. Analyze telemetry and system data to drive improvements in throughput, latency, and resilience. Work cross-functionally with platform, product, and infrastructure teams to deliver seamless storage capabilities across the stack. Share your knowledge and mentor other engineers on best practices in building distributed, high-performance systems. Who You Are: Bachelor's, Master's, or PhD degree in Computer Science, Engineering, or a related field. 8-10+ years of experience working in storage systems engineering or infrastructure. Strong hands-on experience with object storage or distributed filesystems in production environments. Experience with one or more storage protocols (e.g. S3, NFS) and file systems such as Ceph, DAOS, or similar. Proficiency in a systems programming language such as Go, C, or Rust. Proficiency leveraging AI tools to augment software development. Familiarity with storage observability tools and telemetry pipelines (e.g., ClickHouse, Prometheus, Grafana). Experience working with cloud-native infrastructure, Kubernetes, and scalable system architectures. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You love to grow and push the boundaries of your expertise You're curious about distributed storage and the demands AI/ML workloads You're an expert in data persistence on physical media, high performance data transfer using RDMA, or resilient distributed systems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $143,000 to $210,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) , CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: . Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C. 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.