P-78 At Databricks, we are passionate about enabling data teams to solve the world's toughest problems, from advancing AI research to powering next-generation applications. We do this by building and operating the world's best data and AI infrastructure platform. Founded by engineers and driven by customer obsession, we embrace the hardest technical challenges, whether it's scaling distributed systems across multiple clouds or delivering reliable, low-latency communication between thousands of services. And we're only getting started. As a Senior Software Engineer on the Application Traffic team , you will design and build the systems that power Databricks' service-to-service communication across thousands of clusters in a multi-cloud environment. You will also help create abstractions that hide networking complexity from product teams, making connectivity, discovery, and reliability seamless by default. The impact you'll have: You'll work across three key areas that define Databricks' networking stack: Ingress Control Plane : Build the control plane for Databricks' global ingress layer. Enable programming of API gateways with static and dynamic endpoints, simplify service onboarding, and make it easy to expose APIs securely across clouds. Service-to-Service Communication : Design scalable mechanisms for service discovery and load balancing across thousands of clusters. Provide networking abstractions so product teams don't need to worry about underlying connectivity details. Overload Protection : Build intelligent rate limiting and admission control systems to protect critical services under high load. Ensure reliability and predictable performance for both customer-facing and internal workloads. What we look for: BS (or higher) in Computer Science or related field 5+ years of experience designing and building large-scale distributed systems Strong proficiency in one or more languages such as Java, Scala, Go, or C++ Experience with service-oriented architectures and large scale distributed systems Familiarity with cloud platforms (AWS, Azure, GCP) and container/orchestration technologies (Kubernetes, Docker) Track record of shipping infrastructure that supports mission-critical workloads at scale Preferred: background in service discovery, DNS, load balancing, Envoy, or related networking systems Pay Range Transparency Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here. Local Pay Range $166,000-$225,000 USD About Databricks Databricks is the Data and AI company. More than 20,000 organizations worldwide - including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 - rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn, X, YouTube, and Instagram. Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here. Our Commitment to Diversity and Inclusion At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics. Compliance If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.
09/22/2026
Full time
P-78 At Databricks, we are passionate about enabling data teams to solve the world's toughest problems, from advancing AI research to powering next-generation applications. We do this by building and operating the world's best data and AI infrastructure platform. Founded by engineers and driven by customer obsession, we embrace the hardest technical challenges, whether it's scaling distributed systems across multiple clouds or delivering reliable, low-latency communication between thousands of services. And we're only getting started. As a Senior Software Engineer on the Application Traffic team , you will design and build the systems that power Databricks' service-to-service communication across thousands of clusters in a multi-cloud environment. You will also help create abstractions that hide networking complexity from product teams, making connectivity, discovery, and reliability seamless by default. The impact you'll have: You'll work across three key areas that define Databricks' networking stack: Ingress Control Plane : Build the control plane for Databricks' global ingress layer. Enable programming of API gateways with static and dynamic endpoints, simplify service onboarding, and make it easy to expose APIs securely across clouds. Service-to-Service Communication : Design scalable mechanisms for service discovery and load balancing across thousands of clusters. Provide networking abstractions so product teams don't need to worry about underlying connectivity details. Overload Protection : Build intelligent rate limiting and admission control systems to protect critical services under high load. Ensure reliability and predictable performance for both customer-facing and internal workloads. What we look for: BS (or higher) in Computer Science or related field 5+ years of experience designing and building large-scale distributed systems Strong proficiency in one or more languages such as Java, Scala, Go, or C++ Experience with service-oriented architectures and large scale distributed systems Familiarity with cloud platforms (AWS, Azure, GCP) and container/orchestration technologies (Kubernetes, Docker) Track record of shipping infrastructure that supports mission-critical workloads at scale Preferred: background in service discovery, DNS, load balancing, Envoy, or related networking systems Pay Range Transparency Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here. Local Pay Range $166,000-$225,000 USD About Databricks Databricks is the Data and AI company. More than 20,000 organizations worldwide - including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 - rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn, X, YouTube, and Instagram. Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here. Our Commitment to Diversity and Inclusion At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics. Compliance If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.
Senior Staff Engineer, AI Compute (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Capital One machine learning platform organization manages our cloud-based enterprise AI+ML system delivering the high-scale developer and runtime environments required to build, orchestrate, and deploy compute and data intensive AI systems across real-time and batch workloads. We are seeking a Senior Distinguished Engineer, a hands-on technical leader passionate about distributed systems, to engineer and scale foundational compute capabilities for our platform. You will use your experience in building large scale, highly available and high performance systems to develop our common compute infrastructure on top of CPU and GPU substrates. Your contributions will power everything from developer notebooks to ML / DL model training, model inference and feature generation pipelines to pre-training and fine tuning Transformer-based models as well as generative AI inference and agentic applications. Your depth of expertise in technologies including Golang and Python programming languages, popular distributed compute frameworks including Spark / Dask / Ray / Flink, container (e.g., Kubernetes) and serverless (e.g., AWS Lambda) runtime environments, and ML+AI workload patterns will provide an amplifying technical element that is paramount to our team's success. In this role, you will : Architect and build control and data plane implementations required to realize a highly available, multi-tenant, large scale and a secure machine learning platform Develop Ray and Spark distributed compute engine solutions to accelerate diverse workloads from LLM pre-training and reinforcement learning to large-scale data processing, while maximizing compute unit economics Engineer systemic improvements for operational excellence including automating KTLO (Keep The Lights On) workflows Direct the technical execution of a diverse project portfolio, collaborating with developers specializing in everything ranging from distributed microservices to running large foundation models Work cross-functionally with product and program management disciplines, and stakeholder and partners across Capital One to help optimize business outcomes while driving towards strong technology solutions Share your passion for staying on top of tech trends, experimenting with and learning new technologies, participating in internal & external technology communities, and leading system design and code review sessions Help elevate the Capital One Distinguished Engineering community and establish yourself as a go-to resource on given technologies and technology-enabled capabilities Lead the way in creating next-generation talent, mentoring internal talent and actively recruiting external talent to bolster the Capital One tech talent pool Capital One is open to hiring a Remote Employee for this opportunity Basic Qualifications: Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, or Java Preferred Qualifications : Master's Degree in Computer Science or a Master's Degree in Software Engineering Hands on experience in the internals of Ray (Actors/GCS/Scheduling) or Spark (Query Optimizer/Memory Management) Experience building platforms that support LLM training, fine-tuning, or high-throughput inference Hands-on experience with AWS-specific compute primitives (EKS, EC2 UltraClusters, Graviton) and cost-optimization strategies History of upstream contributions to major distributed systems projects Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Distinguished AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Distinguished AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Distinguished AI Engineer New York, NY: $343,400 - $392,000 for Sr. Distinguished AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Distinguished AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/21/2026
Full time
Senior Staff Engineer, AI Compute (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Capital One machine learning platform organization manages our cloud-based enterprise AI+ML system delivering the high-scale developer and runtime environments required to build, orchestrate, and deploy compute and data intensive AI systems across real-time and batch workloads. We are seeking a Senior Distinguished Engineer, a hands-on technical leader passionate about distributed systems, to engineer and scale foundational compute capabilities for our platform. You will use your experience in building large scale, highly available and high performance systems to develop our common compute infrastructure on top of CPU and GPU substrates. Your contributions will power everything from developer notebooks to ML / DL model training, model inference and feature generation pipelines to pre-training and fine tuning Transformer-based models as well as generative AI inference and agentic applications. Your depth of expertise in technologies including Golang and Python programming languages, popular distributed compute frameworks including Spark / Dask / Ray / Flink, container (e.g., Kubernetes) and serverless (e.g., AWS Lambda) runtime environments, and ML+AI workload patterns will provide an amplifying technical element that is paramount to our team's success. In this role, you will : Architect and build control and data plane implementations required to realize a highly available, multi-tenant, large scale and a secure machine learning platform Develop Ray and Spark distributed compute engine solutions to accelerate diverse workloads from LLM pre-training and reinforcement learning to large-scale data processing, while maximizing compute unit economics Engineer systemic improvements for operational excellence including automating KTLO (Keep The Lights On) workflows Direct the technical execution of a diverse project portfolio, collaborating with developers specializing in everything ranging from distributed microservices to running large foundation models Work cross-functionally with product and program management disciplines, and stakeholder and partners across Capital One to help optimize business outcomes while driving towards strong technology solutions Share your passion for staying on top of tech trends, experimenting with and learning new technologies, participating in internal & external technology communities, and leading system design and code review sessions Help elevate the Capital One Distinguished Engineering community and establish yourself as a go-to resource on given technologies and technology-enabled capabilities Lead the way in creating next-generation talent, mentoring internal talent and actively recruiting external talent to bolster the Capital One tech talent pool Capital One is open to hiring a Remote Employee for this opportunity Basic Qualifications: Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, or Java Preferred Qualifications : Master's Degree in Computer Science or a Master's Degree in Software Engineering Hands on experience in the internals of Ray (Actors/GCS/Scheduling) or Spark (Query Optimizer/Memory Management) Experience building platforms that support LLM training, fine-tuning, or high-throughput inference Hands-on experience with AWS-specific compute primitives (EKS, EC2 UltraClusters, Graviton) and cost-optimization strategies History of upstream contributions to major distributed systems projects Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Distinguished AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Distinguished AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Distinguished AI Engineer New York, NY: $343,400 - $392,000 for Sr. Distinguished AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Distinguished AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Distinguished AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
This job is with Warner Bros. Discovery, an inclusive employer and a member of myGwork - the largest global platform for the LGBTQ+ business community. Please do not contact the recruiter directly. Welcome to Warner Bros. Discovery the stuff dreams are made of. Who We Are When we say, "the stuff dreams are made of," we're not just referring to the world of wizards, dragons and superheroes, or even to the wonders of Planet Earth. Behind WBD's vast portfolio of iconic content and beloved brands, are the storytellers bringing our characters to life, the creators bringing them to your living rooms and the dreamers creating what's next From brilliant creatives, to technology trailblazers, across the globe, WBD offers career defining opportunities, thoughtfully curated benefits, and the tools to explore and grow into your best selves. Here you are supported, here you are celebrated, here you can thrive. We are the now and the next. The power behind the people building the future. We are born from the spirit of innovation. We are created from the idea that people around the world want more, need more, deserve more. We are the home of the global digital revolution. We are CNN. To see what it's like to work at CNN, on Instagram and X ! About the Team With deep domain expertise, advanced technical capabilities, and a proven track record of successful collaborations, the AI Enablement & Machine Learning team at CNN is accelerating our digital transformation through strategic applications of machine learning and AI technologies. The ML Platform group within ML Foundations builds and maintains the infrastructure, deployment tooling, and observability that enable CNN's Machine Learning and AI Systems teams to move from prototype to production with velocity and confidence. We support diverse model architectures - two-tower, bandits, LLM- based systems, and traditional ML - across rapid experimentation and scaled production deployment. Our vision is that CNN's ML and AI teams operate with velocity and confidence, supported by infrastructure that handles everything from rapid experimentation to scaled production deployment across diverse model types and use cases. Your New Role As a Staff Software Engineer on ML Platform, you will work across teams to design, build, and operate the infrastructure foundations that power model training, serving, experimentation, and observability for our Machine Learning and AI Systems teams. You will partner with ML engineers, data engineers, and AI Systems engineers to understand production needs, build reliable infrastructure, and deliver tooling that accelerates the team. Key challenges you will tackle: Outerbounds Migration: Own end-to-end migration of CNN's ML training and orchestration workloads to Outerbounds-managed Metaflow, with zero production disruption and a clear path to self-service for ML practitioners. Observability and Cost Attribution: Establish comprehensive observability across all ML products, AI applications, and shared infrastructure - including the monitoring, alerting, and diagnostic tooling engineers need to operate production systems reliably, and cost attribution that lets us understand and govern ML/AI spend by team, product, and use case. Model and Application Registry: Design and operate a unified registry and versioning system for ML models and AI applications, providing lineage, reproducibility, and a clean handoff between development and production. Feature Resolution Framework: Evolve our existing feature resolution framework from a loosely coupled set of pipelines into a real platform - one that lets ML and AI Systems engineers configure content types and events to intercept, register feature-generation APIs, and land resolved features as durable data products in the feature store. The features this framework produces - ML- and LLM-generated alike - power everything from analytics to training to inference to user-facing rendering. What You'll Do Design and own infrastructure, deployment tooling, and developer experience for ML and AI Systems teams Lead architectural decisions across orchestration, serving, observability, and experimentation infrastructure Build self-service tooling that lets ML practitioners move from prototype to production without platform team dependencies Establish engineering standards for ML/AI infrastructure, including reliability, cost governance, and operational excellence Review designs and code, mentor engineers, and lead cross-team initiatives Partner with Data Platform on infrastructure coordination and data access patterns Communicate effectively across audiences - technical documentation, design reviews, and stakeholder interactions The Essentials 8+ years building production infrastructure or platform systems, with a Bachelor's degree in Computer Science, Information Technology, or a related technical field (or 6+ years with a Master's degree) Deep expertise in distributed systems, with a track record of shipping highly available, low-latency infrastructure Strong proficiency in Python and at least one of Go, Java, or C++ Expertise with cloud infrastructure and IaC, especially AWS and Terraform Experience with ML or data infrastructure - orchestration, serving, deployment tooling, observability, or experimentation frameworks Proven track record of leading complex platform projects from concept to production - knowing when to own decisions, when to rally the right people for alignment, and when to escalate Collaborative mindset, understanding that great platform work depends on deep partnership with the teams you serve A passion for helping CNN's engineering organization grow through mentorship, talent acquisition, and professional development The Nice to Haves Experience with Metaflow, SageMaker, or comparable ML orchestration platforms Experience with model registries, feature stores, or experimentation frameworks Background in cost governance, FinOps, or multi-tenant infrastructure Practical experience supporting LLM-based or GenAI production systems Prior experience working closely with machine learning engineers How We Get Things Done This last bit is probably the most important! Here at WBD, our guiding principles are the core values by which we operate and are central to how we get things done. You can find them at along with some insights from the team on what they mean and how they show up in their day to day. We hope they resonate with you and look forward to discussing them during your interview. Championing Inclusion at WBD Warner Bros. Discovery embraces the opportunity to build a workforce that reflects a wide array of perspectives, backgrounds and experiences. Being an equal opportunity employer means that we take seriously our responsibility to consider qualified candidates on the basis of merit, without regard to race, color, religion, national origin, gender, sexual orientation, gender identity or expression, age, mental or physical disability, and genetic information, marital status, citizenship status, military status, protected veteran status or any other category protected by law. If you're a qualified candidate with a disability and you require adjustments or accommodations during the job application and/or recruitment process, please visit our accessibility page for instructions to submit your request.
09/21/2026
Full time
This job is with Warner Bros. Discovery, an inclusive employer and a member of myGwork - the largest global platform for the LGBTQ+ business community. Please do not contact the recruiter directly. Welcome to Warner Bros. Discovery the stuff dreams are made of. Who We Are When we say, "the stuff dreams are made of," we're not just referring to the world of wizards, dragons and superheroes, or even to the wonders of Planet Earth. Behind WBD's vast portfolio of iconic content and beloved brands, are the storytellers bringing our characters to life, the creators bringing them to your living rooms and the dreamers creating what's next From brilliant creatives, to technology trailblazers, across the globe, WBD offers career defining opportunities, thoughtfully curated benefits, and the tools to explore and grow into your best selves. Here you are supported, here you are celebrated, here you can thrive. We are the now and the next. The power behind the people building the future. We are born from the spirit of innovation. We are created from the idea that people around the world want more, need more, deserve more. We are the home of the global digital revolution. We are CNN. To see what it's like to work at CNN, on Instagram and X ! About the Team With deep domain expertise, advanced technical capabilities, and a proven track record of successful collaborations, the AI Enablement & Machine Learning team at CNN is accelerating our digital transformation through strategic applications of machine learning and AI technologies. The ML Platform group within ML Foundations builds and maintains the infrastructure, deployment tooling, and observability that enable CNN's Machine Learning and AI Systems teams to move from prototype to production with velocity and confidence. We support diverse model architectures - two-tower, bandits, LLM- based systems, and traditional ML - across rapid experimentation and scaled production deployment. Our vision is that CNN's ML and AI teams operate with velocity and confidence, supported by infrastructure that handles everything from rapid experimentation to scaled production deployment across diverse model types and use cases. Your New Role As a Staff Software Engineer on ML Platform, you will work across teams to design, build, and operate the infrastructure foundations that power model training, serving, experimentation, and observability for our Machine Learning and AI Systems teams. You will partner with ML engineers, data engineers, and AI Systems engineers to understand production needs, build reliable infrastructure, and deliver tooling that accelerates the team. Key challenges you will tackle: Outerbounds Migration: Own end-to-end migration of CNN's ML training and orchestration workloads to Outerbounds-managed Metaflow, with zero production disruption and a clear path to self-service for ML practitioners. Observability and Cost Attribution: Establish comprehensive observability across all ML products, AI applications, and shared infrastructure - including the monitoring, alerting, and diagnostic tooling engineers need to operate production systems reliably, and cost attribution that lets us understand and govern ML/AI spend by team, product, and use case. Model and Application Registry: Design and operate a unified registry and versioning system for ML models and AI applications, providing lineage, reproducibility, and a clean handoff between development and production. Feature Resolution Framework: Evolve our existing feature resolution framework from a loosely coupled set of pipelines into a real platform - one that lets ML and AI Systems engineers configure content types and events to intercept, register feature-generation APIs, and land resolved features as durable data products in the feature store. The features this framework produces - ML- and LLM-generated alike - power everything from analytics to training to inference to user-facing rendering. What You'll Do Design and own infrastructure, deployment tooling, and developer experience for ML and AI Systems teams Lead architectural decisions across orchestration, serving, observability, and experimentation infrastructure Build self-service tooling that lets ML practitioners move from prototype to production without platform team dependencies Establish engineering standards for ML/AI infrastructure, including reliability, cost governance, and operational excellence Review designs and code, mentor engineers, and lead cross-team initiatives Partner with Data Platform on infrastructure coordination and data access patterns Communicate effectively across audiences - technical documentation, design reviews, and stakeholder interactions The Essentials 8+ years building production infrastructure or platform systems, with a Bachelor's degree in Computer Science, Information Technology, or a related technical field (or 6+ years with a Master's degree) Deep expertise in distributed systems, with a track record of shipping highly available, low-latency infrastructure Strong proficiency in Python and at least one of Go, Java, or C++ Expertise with cloud infrastructure and IaC, especially AWS and Terraform Experience with ML or data infrastructure - orchestration, serving, deployment tooling, observability, or experimentation frameworks Proven track record of leading complex platform projects from concept to production - knowing when to own decisions, when to rally the right people for alignment, and when to escalate Collaborative mindset, understanding that great platform work depends on deep partnership with the teams you serve A passion for helping CNN's engineering organization grow through mentorship, talent acquisition, and professional development The Nice to Haves Experience with Metaflow, SageMaker, or comparable ML orchestration platforms Experience with model registries, feature stores, or experimentation frameworks Background in cost governance, FinOps, or multi-tenant infrastructure Practical experience supporting LLM-based or GenAI production systems Prior experience working closely with machine learning engineers How We Get Things Done This last bit is probably the most important! Here at WBD, our guiding principles are the core values by which we operate and are central to how we get things done. You can find them at along with some insights from the team on what they mean and how they show up in their day to day. We hope they resonate with you and look forward to discussing them during your interview. Championing Inclusion at WBD Warner Bros. Discovery embraces the opportunity to build a workforce that reflects a wide array of perspectives, backgrounds and experiences. Being an equal opportunity employer means that we take seriously our responsibility to consider qualified candidates on the basis of merit, without regard to race, color, religion, national origin, gender, sexual orientation, gender identity or expression, age, mental or physical disability, and genetic information, marital status, citizenship status, military status, protected veteran status or any other category protected by law. If you're a qualified candidate with a disability and you require adjustments or accommodations during the job application and/or recruitment process, please visit our accessibility page for instructions to submit your request.
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration. Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models. This is a ground-up effort with rapidly evolving hardware and software. We are looking for an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack. Key job responsibilities - Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference. - Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware. - Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism. - Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets. - Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up. - Own features end-to-end: from design through implementation, testing, and integration into the broader software stack. - Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions. BASIC QUALIFICATIONS - Bachelor's degree or equivalent - 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Knowledge of computer architecture, operating systems, and parallel computing - Knowledge of Linux fundamentals - Strong proficiency in C/C++ - Experience developing compute kernels for GPUs, DSPs, or custom accelerators - Proven track record of owning and delivering complex software features end-to-end PREFERRED QUALIFICATIONS - Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming - Experience with hardware simulation environments and model validation workflows - Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually
09/20/2026
Full time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration. Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models. This is a ground-up effort with rapidly evolving hardware and software. We are looking for an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack. Key job responsibilities - Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference. - Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware. - Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism. - Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets. - Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up. - Own features end-to-end: from design through implementation, testing, and integration into the broader software stack. - Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions. BASIC QUALIFICATIONS - Bachelor's degree or equivalent - 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Knowledge of computer architecture, operating systems, and parallel computing - Knowledge of Linux fundamentals - Strong proficiency in C/C++ - Experience developing compute kernels for GPUs, DSPs, or custom accelerators - Proven track record of owning and delivering complex software features end-to-end PREFERRED QUALIFICATIONS - Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming - Experience with hardware simulation environments and model validation workflows - Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, CA, Cupertino - 165 600.00 USD annually
Job Description Job Description About Etched Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history. Job Summary Etched's infrastructure spans some of the most sensitive compute environments in the industry: bare-metal HPC clusters running proprietary ASIC workloads, hybrid on-prem/cloud deployments, and internal toolchains that house irreplaceable chip design IP. As we scale from early silicon to production, securing these environments is foundational - not an afterthought. As our first dedicated Network Security Engineer, you will own the design and implementation of Etched's network security posture end to end. You'll work alongside the infrastructure team to harden our physical and virtual networks, enforce least-privilege access to chip design environments, and build the detection and response capabilities that keep our most sensitive assets safe. This is a high-ownership role for someone who wants to shape security architecture at a company building the compute infrastructure for the next decade of AI - not maintain someone else's stack. Key Responsibilities Design and implement a zero-trust network architecture across on-prem datacenters, multiple office locations, and multi-cloud platforms, including secure remote access that eliminates VPN sprawl without sacrificing engineer usability and speed Define and enforce network segmentation policies that isolate sensitive ASIC development workflows from general infrastructure, customer access, validation labs, and manufacturing infrastructure Balancing prevention and detection, deploy, tune, and operate NDR, IDS/IPS, and next-generation firewalls across our physical and virtual network fabric; build automation to continuously assess and enforce firewall rules, ACLs, and routing policies - treating network security configuration as code Integrate and operate EDR/XDR, MDM/MAM, SASE, and CASB tooling in partnership with end-user and IT teams, enforcing unified DLP policies and device compliance posture across endpoint, cloud, and network control planes to eliminate data exfiltration risk Own our vulnerability management process for network-layer exposure: scanning, prioritization, and remediation tracking in partnership with infrastructure engineers Lead incident response for network-layer security events: detection, containment, root-cause analysis, and post-incident hardening Partner with legal, compliance, and leadership to support regulatory requirements and customer security reviews as they arise Architect and deploy network segmentation for our HPC clusters, isolating EDA tool traffic, ASIC simulation workloads, and CI pipelines from each other and from the corporate network Architect and deploy a ZTNA-based corporate network that eliminates VPN sprawl and ensures end-user devices maintain a consistent security posture and seamless access to sensitive development environments - whether engineers are on-site, remote, or traveling - replacing location-dependent trust with continuous identity and device health verification Design and implement a scalable NDR pipeline that ingests flow data across bare-metal switches and cloud VPCs, feeds a centralized SIEM, and generates actionable alerts with low false-positive rates Develop runbooks and automated playbooks for the highest-probability incident scenarios - credential compromise, lateral movement, and exfiltration from IP-sensitive environments Integrate EDR/XDR telemetry with SASE enforcement and CASB inline controls to build a unified DLP detection and response pipeline spanning endpoints, cloud SaaS, and the corporate network Partner with end-user and IT teams to roll out MDM/MAM policies that containerize sensitive IP on engineer devices and enforce compliance-based conditional access across managed and unmanaged environments You may be a good fit if you have (Must-have qualifications) Bring deep, broad networking expertise - from low-level packet analysis and firewall log forensics to BGP configuration, multi-cloud networking, and CASB/SASE integration across a diverse SaaS landscape Have hands-on experience with the Fortinet ecosystem - firewalls, FortiSASE, FortiAPs, and switches - and are comfortable with Arista switch platforms, including configuration, EOS automation, and integration into a broader security architecture Treat security as an engineering discipline: you write code and automation rather than relying on point-and-click tooling, version-control your configurations, and develop intent-driven network automation Have experience securing high-value compute environments - datacenters, HPC clusters, semiconductor design environments, or similar settings where the cost of a breach is extremely high Have deployed and integrated EDR/XDR, MDM/MAM, SASE, and CASB tooling, and understand how to stitch them together into a unified DLP and access control framework that spans endpoints, cloud, and the network Have built or operated ZTNA-based access models and understand how to enforce consistent security posture across on-site, remote, and traveling users without degrading the experience for engineers Are comfortable owning your domain with minimal oversight: you can independently scope a project, identify the right tooling, and drive it to completion Have strong Linux fundamentals and understand how OS-level networking (iptables/nftables, network namespaces, eBPF) interacts with physical and virtual network security controls Have built or operated network security monitoring at scale - you know the difference between a good alert and noise, and you can architect a detection pipeline that surfaces real signal Can communicate risk clearly to both technical peers and non-technical leadership, and can translate security requirements into actionable infrastructure changes Strong candidates may also have experience with (Nice-to-have qualifications) Experience with EDA environments or semiconductor IP security Familiarity with cloud-native network security controls on AWS, GCP, or Azure (security groups, VPC flow logs, cloud firewalls, CSPM) Background in or exposure to NIST, SOC 2, or ISO 27001 frameworks Experience with eBPF-based network observability and security tooling Benefits Medical, dental, and vision packages with generous premium coverage $500 per month credit for waiving medical benefits Housing subsidy of $2k per month for those living within walking distance of the office Relocation support for those moving to San Jose (Santana Row) Various wellness benefits covering fitness, mental health, and more Daily lunch and dinner in our office Unlimited compute budget subject to ROI justification How we're different Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors. We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed. Compensation Range: $175K - $275K
09/15/2026
Full time
Job Description Job Description About Etched Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history. Job Summary Etched's infrastructure spans some of the most sensitive compute environments in the industry: bare-metal HPC clusters running proprietary ASIC workloads, hybrid on-prem/cloud deployments, and internal toolchains that house irreplaceable chip design IP. As we scale from early silicon to production, securing these environments is foundational - not an afterthought. As our first dedicated Network Security Engineer, you will own the design and implementation of Etched's network security posture end to end. You'll work alongside the infrastructure team to harden our physical and virtual networks, enforce least-privilege access to chip design environments, and build the detection and response capabilities that keep our most sensitive assets safe. This is a high-ownership role for someone who wants to shape security architecture at a company building the compute infrastructure for the next decade of AI - not maintain someone else's stack. Key Responsibilities Design and implement a zero-trust network architecture across on-prem datacenters, multiple office locations, and multi-cloud platforms, including secure remote access that eliminates VPN sprawl without sacrificing engineer usability and speed Define and enforce network segmentation policies that isolate sensitive ASIC development workflows from general infrastructure, customer access, validation labs, and manufacturing infrastructure Balancing prevention and detection, deploy, tune, and operate NDR, IDS/IPS, and next-generation firewalls across our physical and virtual network fabric; build automation to continuously assess and enforce firewall rules, ACLs, and routing policies - treating network security configuration as code Integrate and operate EDR/XDR, MDM/MAM, SASE, and CASB tooling in partnership with end-user and IT teams, enforcing unified DLP policies and device compliance posture across endpoint, cloud, and network control planes to eliminate data exfiltration risk Own our vulnerability management process for network-layer exposure: scanning, prioritization, and remediation tracking in partnership with infrastructure engineers Lead incident response for network-layer security events: detection, containment, root-cause analysis, and post-incident hardening Partner with legal, compliance, and leadership to support regulatory requirements and customer security reviews as they arise Architect and deploy network segmentation for our HPC clusters, isolating EDA tool traffic, ASIC simulation workloads, and CI pipelines from each other and from the corporate network Architect and deploy a ZTNA-based corporate network that eliminates VPN sprawl and ensures end-user devices maintain a consistent security posture and seamless access to sensitive development environments - whether engineers are on-site, remote, or traveling - replacing location-dependent trust with continuous identity and device health verification Design and implement a scalable NDR pipeline that ingests flow data across bare-metal switches and cloud VPCs, feeds a centralized SIEM, and generates actionable alerts with low false-positive rates Develop runbooks and automated playbooks for the highest-probability incident scenarios - credential compromise, lateral movement, and exfiltration from IP-sensitive environments Integrate EDR/XDR telemetry with SASE enforcement and CASB inline controls to build a unified DLP detection and response pipeline spanning endpoints, cloud SaaS, and the corporate network Partner with end-user and IT teams to roll out MDM/MAM policies that containerize sensitive IP on engineer devices and enforce compliance-based conditional access across managed and unmanaged environments You may be a good fit if you have (Must-have qualifications) Bring deep, broad networking expertise - from low-level packet analysis and firewall log forensics to BGP configuration, multi-cloud networking, and CASB/SASE integration across a diverse SaaS landscape Have hands-on experience with the Fortinet ecosystem - firewalls, FortiSASE, FortiAPs, and switches - and are comfortable with Arista switch platforms, including configuration, EOS automation, and integration into a broader security architecture Treat security as an engineering discipline: you write code and automation rather than relying on point-and-click tooling, version-control your configurations, and develop intent-driven network automation Have experience securing high-value compute environments - datacenters, HPC clusters, semiconductor design environments, or similar settings where the cost of a breach is extremely high Have deployed and integrated EDR/XDR, MDM/MAM, SASE, and CASB tooling, and understand how to stitch them together into a unified DLP and access control framework that spans endpoints, cloud, and the network Have built or operated ZTNA-based access models and understand how to enforce consistent security posture across on-site, remote, and traveling users without degrading the experience for engineers Are comfortable owning your domain with minimal oversight: you can independently scope a project, identify the right tooling, and drive it to completion Have strong Linux fundamentals and understand how OS-level networking (iptables/nftables, network namespaces, eBPF) interacts with physical and virtual network security controls Have built or operated network security monitoring at scale - you know the difference between a good alert and noise, and you can architect a detection pipeline that surfaces real signal Can communicate risk clearly to both technical peers and non-technical leadership, and can translate security requirements into actionable infrastructure changes Strong candidates may also have experience with (Nice-to-have qualifications) Experience with EDA environments or semiconductor IP security Familiarity with cloud-native network security controls on AWS, GCP, or Azure (security groups, VPC flow logs, cloud firewalls, CSPM) Background in or exposure to NIST, SOC 2, or ISO 27001 frameworks Experience with eBPF-based network observability and security tooling Benefits Medical, dental, and vision packages with generous premium coverage $500 per month credit for waiving medical benefits Housing subsidy of $2k per month for those living within walking distance of the office Relocation support for those moving to San Jose (Santana Row) Various wellness benefits covering fitness, mental health, and more Daily lunch and dinner in our office Unlimited compute budget subject to ROI justification How we're different Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors. We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed. Compensation Range: $175K - $275K
Architect, build, and continuously evolve a truly declarative, self-healing, cloud-native, Kubernetes-native, and highly observable enterprise-grade platform. This strategic leadership role defines and drives the fundamental architecture of our technology stack, including: Go-based microservices architecture, ensuring high throughput and resilience; next-generation CI/CD pipelines utilizing GitHub Actions and Azure DevOps, engineered for speed, reliability, governance, and release integrity; infrastructure-as-Code (IaC) automation using Terraform and Ansible to deliver consistent, secure, and repeatable environments at scale; world-class Observability strategy, integrating metrics, logging, and tracing, and intelligent alerting to enable proactive operations and optimization and governance of MongoDB data persistence and distributed systems, ensuring performance, durability, and operational efficiency. 1. CI/CD Strategy & Governance (The Quality Pipeline) Responsibility & Authority (R&A) for the end-to-end architecture, standardization, and all YAML-based delivery pipelines (build, test, security scanning, artifact management). Establish and enforce mandatory, non-negotiable quality gates (unit, integration, security, performance) at every stage of the pipeline to ensure "shift-left" quality. Secure and harden the CI/CD ecosystem against tampering and supply chain risks through strong access controls, artifact integrity, and policy enforcement. 2. Deployment Excellence & Safety (Production Integrity) Define, standardize, and govern safe, high-integrity deployment strategies (e.g., Canary Releases, Rolling Updates, Blue/Green, utilization of Feature Flags). Implement robust, fast, and fully automated rollback mechanisms, ensuring absolute production safety and minimizing Mean Time To Recovery (MTTR). Continuously reduce deployment risk and complexity through automation, standardization, and release discipline. 3. Production Reliability Engineering (SRE & Resilience) Establish, evangelize, and enforce enterprise-grade standards for Stability, Performance, Scalability, Observability, and Disaster Recovery (DR). Develop, monitor, and report on key Service Level Objectives (SLOs) and Service Level Agreements (SLAs). Lead and mature the incident response lifecycle, focusing on root cause analysis and preventative, long-term remediation. The focus is on achieving production resilience. 4. AI-Accelerated DevOps (Responsible Innovation) Champion the responsible and ethical integration of generative AI tools to accelerate repeatable engineering tasks such as IaC generation, test case development, and documentation creation. Implement prompt engineering practices to maximize AI effectiveness and reliability. Establish and enforce rigorous verification and validation processes for all AI-generated artifacts. Increased velocity must never compromise engineering correctness, security, or compliance. 5. Platform Architecture & Optimization (Scalability & Performance) Oversee the continuous evolution of the scalable, resilient architecture for the foundational Kubernetes control plane. Drive the design of resilient Go microservices and REST APIs, emphasizing security and performance. Provide expertise in optimizing MongoDB performance, indexing strategies, and distributed systems design to ensure 24/7 availability and data integrity. Education and Experience Requirements BS degree in Computer Science, Computer Engineering or related plus 10 years of experience as a Engineering Manager, Software Engineer, or related. Special Skills Requirements Requires 10 years of experience in software engineering, DevOps, or Site Reliability Engineering (SRE). Requires 5 years of programming experience in Python, Go, Java, or a comparable enterprise-grade language; in operating, scaling, and troubleshooting Kubernetes within large, complex production environments; in MongoDB schema design, query optimization, and performance tuning for high volume, mission-critical environments; designing, governing, and securing modern CI/CD pipelines (e.g., Jenkins, GitHub Actions, Azure DevOps) with strong emphasis on automation and compliance controls; and Infrastructure-as-Code (Terraform and/or Ansible) and with one or more major cloud platforms (AWS, Azure, or GCP). Requires 3 years of direct technical leadership experience, leading teams supporting high-scale distributed systems. Please copy and paste your resume in the email body (do not send attachments, we cannot open them) and email it to candidates at (link removed) with reference in the subject line. Thank you.
09/15/2026
Architect, build, and continuously evolve a truly declarative, self-healing, cloud-native, Kubernetes-native, and highly observable enterprise-grade platform. This strategic leadership role defines and drives the fundamental architecture of our technology stack, including: Go-based microservices architecture, ensuring high throughput and resilience; next-generation CI/CD pipelines utilizing GitHub Actions and Azure DevOps, engineered for speed, reliability, governance, and release integrity; infrastructure-as-Code (IaC) automation using Terraform and Ansible to deliver consistent, secure, and repeatable environments at scale; world-class Observability strategy, integrating metrics, logging, and tracing, and intelligent alerting to enable proactive operations and optimization and governance of MongoDB data persistence and distributed systems, ensuring performance, durability, and operational efficiency. 1. CI/CD Strategy & Governance (The Quality Pipeline) Responsibility & Authority (R&A) for the end-to-end architecture, standardization, and all YAML-based delivery pipelines (build, test, security scanning, artifact management). Establish and enforce mandatory, non-negotiable quality gates (unit, integration, security, performance) at every stage of the pipeline to ensure "shift-left" quality. Secure and harden the CI/CD ecosystem against tampering and supply chain risks through strong access controls, artifact integrity, and policy enforcement. 2. Deployment Excellence & Safety (Production Integrity) Define, standardize, and govern safe, high-integrity deployment strategies (e.g., Canary Releases, Rolling Updates, Blue/Green, utilization of Feature Flags). Implement robust, fast, and fully automated rollback mechanisms, ensuring absolute production safety and minimizing Mean Time To Recovery (MTTR). Continuously reduce deployment risk and complexity through automation, standardization, and release discipline. 3. Production Reliability Engineering (SRE & Resilience) Establish, evangelize, and enforce enterprise-grade standards for Stability, Performance, Scalability, Observability, and Disaster Recovery (DR). Develop, monitor, and report on key Service Level Objectives (SLOs) and Service Level Agreements (SLAs). Lead and mature the incident response lifecycle, focusing on root cause analysis and preventative, long-term remediation. The focus is on achieving production resilience. 4. AI-Accelerated DevOps (Responsible Innovation) Champion the responsible and ethical integration of generative AI tools to accelerate repeatable engineering tasks such as IaC generation, test case development, and documentation creation. Implement prompt engineering practices to maximize AI effectiveness and reliability. Establish and enforce rigorous verification and validation processes for all AI-generated artifacts. Increased velocity must never compromise engineering correctness, security, or compliance. 5. Platform Architecture & Optimization (Scalability & Performance) Oversee the continuous evolution of the scalable, resilient architecture for the foundational Kubernetes control plane. Drive the design of resilient Go microservices and REST APIs, emphasizing security and performance. Provide expertise in optimizing MongoDB performance, indexing strategies, and distributed systems design to ensure 24/7 availability and data integrity. Education and Experience Requirements BS degree in Computer Science, Computer Engineering or related plus 10 years of experience as a Engineering Manager, Software Engineer, or related. Special Skills Requirements Requires 10 years of experience in software engineering, DevOps, or Site Reliability Engineering (SRE). Requires 5 years of programming experience in Python, Go, Java, or a comparable enterprise-grade language; in operating, scaling, and troubleshooting Kubernetes within large, complex production environments; in MongoDB schema design, query optimization, and performance tuning for high volume, mission-critical environments; designing, governing, and securing modern CI/CD pipelines (e.g., Jenkins, GitHub Actions, Azure DevOps) with strong emphasis on automation and compliance controls; and Infrastructure-as-Code (Terraform and/or Ansible) and with one or more major cloud platforms (AWS, Azure, or GCP). Requires 3 years of direct technical leadership experience, leading teams supporting high-scale distributed systems. Please copy and paste your resume in the email body (do not send attachments, we cannot open them) and email it to candidates at (link removed) with reference in the subject line. Thank you.