Senior / Lead .NET Software Engineer / SRE DETAILS Location: Arlington, TX 76014 (hybrid onsite 2-days per week) Openings: 2 (1 Senior .NET / 1 Lead .NET) Position Type: 6M C2H Hourly / Salary: to $160K+ (based on experience level) JOB SUMMARY Vaco is currently seeking a Senior .NET Software Engineer / SRE for a 6M C2H opportunity that is located in Arlington, TX 76014 (hybrid onsite 2-days per week). About the Project: This is an ongoing multi-year digital transformation initiative (now roughly 2+ years in) focused on modernizing financial services platforms, making them more scalable, reliable, automated, and cloud-native. The .NET Engineer / SRE will be building and extending tools / frameworks for SRE practices, including automation scripts, custom monitoring, CI/CD enhancements, reliability tooling, etc. The .NET Engineer / SRE will help bridge the gap between development teams (writing .NET) and SREs (ensuring reliability at scale), especially within this strategic / tool building-focused SRE practice. This is not a traditional .NET Developer (no owning user-facing APIs or business logic for financial services), but rather applying .NET skills to SRE tasks, including building custom tools for automation, reliability, and release engineering, creating GITHub Copilot context files to embed SRE practices (security / performance checks) early in development cycles, debugging / refactoring code in production environments, performing root-cause analysis, and iterating post-deployment. Additionally, there's no traditional on-call production monitoring / support, whereas the focus will be on design, architecture, automation, and enabling scalable DevOps / SRE practices for development teams, rather than day-to-day firefighting. Ideal Senior Candidate:The ideal candidate will come with a Developer-First Mindset for full SDLC, Architecture Ownership, and Problem Solving (build the tool if it does not exist approach). Having .NET fluency enables strategic SRE in their .NET / Azure ecosystem, bridging gaps between development and operations without siloed roles. The Senior Engineer MUST be able to walk through their most recent projects, in depth, describing complex features, debugging, and development frameworks (why you'd choose one over the other for particular tasks, etc.). Ideal LEAD Candidate (leading teams of 5+): The ideal candidate will come with a Developer-First Mindset for full SDLC, Architecture Ownership, and Problem Solving (build the tool if it does not exist approach). Having .NET fluency enables strategic SRE in their .NET / Azure ecosystem, bridging gaps between development and operations without siloed roles. The Lead Engineer MUST be able to walk through their most recent projects, in depth, describing complex features, debugging, and development frameworks (why you'd choose one over the other for particular tasks, etc.). JOB REQUIREMETS .NET Core Development (hands-on core expertise) - C# / .NET Core Development Building / Maintaining Production-Grade Applications / APIs / Services Cloud Modernization - Migrating Legacy .NET Applications to Cloud-Native Azure Architectures Azure-Native Application Development - Designing / Deploying / Supporting .NET Applications Directly in Azure with Real Production Workloads (Beyond Portal-Based "Click-Ops") AKS / Containerization - Containerizing .NET Applications (Docker Multi-Stage Builds) Deploying / Managing within AKS (Deployments / Services / ConfigMaps / Secrets / Helm / Node Pools / Autoscaling / Ingress / Cluster-Level Troubleshooting) CI/CD / Azure DevOps - Building / Maintaining YAML-Based CI/CD Pipelines (Automated Builds / Testing / Security Scans / Docker Image Creation / Deployments to AKS / App Services) Azure Platform Services Integration - Integrating Core Azure Services into .NET Applications (App Services / Azure Functions / Service Bus / Event Grid / Azure Monitor / Application Insights With Custom Telemetry / Distributed Tracing / Metrics) IaC / Terraform - Developing / Managing Terraform-Based Infrastructure (AKS / VNETs / Private Endpoints / Service Bus / SQL Database / Storage) Aligned with Application Development Identity / Security (Azure) - Implementing Secure Patterns (Managed Identities / Entra ID / RBAC / Secure Connectivity / Polly Retry Policies / Private Endpoints / VNET Integration) within .NET Services Observability / SRE Practices - Applying SRE Principles within Development (Application Insights / Log Analytics / Alerting / Performance Monitoring / Production Troubleshooting) Database Design / Optimization - Oracle / MS SQL Server / NoSQL (CosmosDB) SQL Scripting (hands-on) Designing / Evolving Database Schemas Performing Query Performance Analysis Indexing to Deliver Scalable / Performant Services Problem Solving / Collaboration - Driving Root Cause Analysis / Debugging for .NET Applications in Azure Collaborating within SRE / Cross-Functional Teams PREFERRED (not required) Azure Governance / Best Practices - Azure Policy / GitHub Copilot Context / Embedding SRE Practices Early in Development Lifecycle Architecture Frameworks (Familiarity) - Azure Well-Architected Framework (Reliability / Operational Excellence) Advanced AKS / Hybrid - Exposure to Advanced AKS Topics (Cluster Autoscaling / Service Mesh / Cost Optimization) / Hybrid Scenarios (Azure Arc) AI / Automation Enablement - Driving Adoption of AI-Powered Tools (GitHub Copilot) to Improve Developer Productivity / Code Quality By submitting to this position, you are agreeing to be included in our talent pool for future hiring for similarly qualified positions. EEO Notice Vaco by Highspring is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of race (including but not limited to traits historically associated with race such as hair texture and hair style), color, sex (includes pregnancy or related conditions), religion or creed, national origin, citizenship, age, disability, status as a veteran, union membership, ethnicity, gender, gender identity, gender expression, sexual orientation, marital status, political affiliation, or any other protected characteristics as required by federal, state or local law. Vaco by Highspring and its parents, affiliates, and subsidiaries are committed to the full inclusion of all qualified individuals. As part of this commitment, Vaco by Highspring and its parents, affiliates, and subsidiaries will ensure that persons with disabilities are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact . Vaco by Highspring also wants all applicants to know their rights that workplace discrimination is illegal. Representation Notice By submitting to this position, you agree that you will be giving Vaco by Highspring the exclusive right to present your as a candidate for the foregoing employment opportunity. You further agree that you have represented information about yourself accurately and have not affirmatively misrepresented your qualifications. You also agree to maintain as confidential, to the fullest extent permitted by law, any information you learn from Vaco by Highspring about the position and you will limit disclosure of information about the position only to the extent necessary to perform any obligations in furtherance of your application. In exchange, Vaco by Highspring agrees to exercise reasonable efforts to represent you through all solicitation, job screening and resume dispersal. For residents of Ontario, Canada: Based on Highspring's discussions with its Client, Highspring's understanding is that this position for employment is a current vacancy (either through Highspring as a contractor or with the client directly). Privacy Notice Vaco by Highspring and its parents, affiliates, and subsidiaries ("we," "our," or "Vaco by Highspring") respects your privacy and are committed to providing transparent notice of our policies. California residents may access Vaco by Highspring HR Notice at Collection for California Applicants and Employees here. Virginia residents may access our state specific policies here. Residents of all other states may access our policies here. Canadian residents may access our policies in English here and in French here. Residents of countries governed by GDPR may access our policies here. Additionally, submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. More details about Vaco by Highspring's use of AI can be found here (). Further assessment of candidates beyond this initial phase will be conducted by recruiters and hiring managers. Vaco by Highspring does not know and cannot opine on if its client's use of AI products in hiring. Pay Transparency Notice Determining compensation for this role (and others) at Vaco by Highspring depends upon a wide array of factors including but not limited to: the individual's skill sets, experience and training; licensure and certification requirements; office location and other geographic considerations; other business and organizational needs. With that said, as required by local law, Vaco by Highspring believes that the following salary range referenced above reasonably estimates the base compensation for an individual hired into this position in geographies that require salary range disclosure. The individual may also be eligible for discretionary bonuses.
09/25/2026
Full time
Senior / Lead .NET Software Engineer / SRE DETAILS Location: Arlington, TX 76014 (hybrid onsite 2-days per week) Openings: 2 (1 Senior .NET / 1 Lead .NET) Position Type: 6M C2H Hourly / Salary: to $160K+ (based on experience level) JOB SUMMARY Vaco is currently seeking a Senior .NET Software Engineer / SRE for a 6M C2H opportunity that is located in Arlington, TX 76014 (hybrid onsite 2-days per week). About the Project: This is an ongoing multi-year digital transformation initiative (now roughly 2+ years in) focused on modernizing financial services platforms, making them more scalable, reliable, automated, and cloud-native. The .NET Engineer / SRE will be building and extending tools / frameworks for SRE practices, including automation scripts, custom monitoring, CI/CD enhancements, reliability tooling, etc. The .NET Engineer / SRE will help bridge the gap between development teams (writing .NET) and SREs (ensuring reliability at scale), especially within this strategic / tool building-focused SRE practice. This is not a traditional .NET Developer (no owning user-facing APIs or business logic for financial services), but rather applying .NET skills to SRE tasks, including building custom tools for automation, reliability, and release engineering, creating GITHub Copilot context files to embed SRE practices (security / performance checks) early in development cycles, debugging / refactoring code in production environments, performing root-cause analysis, and iterating post-deployment. Additionally, there's no traditional on-call production monitoring / support, whereas the focus will be on design, architecture, automation, and enabling scalable DevOps / SRE practices for development teams, rather than day-to-day firefighting. Ideal Senior Candidate:The ideal candidate will come with a Developer-First Mindset for full SDLC, Architecture Ownership, and Problem Solving (build the tool if it does not exist approach). Having .NET fluency enables strategic SRE in their .NET / Azure ecosystem, bridging gaps between development and operations without siloed roles. The Senior Engineer MUST be able to walk through their most recent projects, in depth, describing complex features, debugging, and development frameworks (why you'd choose one over the other for particular tasks, etc.). Ideal LEAD Candidate (leading teams of 5+): The ideal candidate will come with a Developer-First Mindset for full SDLC, Architecture Ownership, and Problem Solving (build the tool if it does not exist approach). Having .NET fluency enables strategic SRE in their .NET / Azure ecosystem, bridging gaps between development and operations without siloed roles. The Lead Engineer MUST be able to walk through their most recent projects, in depth, describing complex features, debugging, and development frameworks (why you'd choose one over the other for particular tasks, etc.). JOB REQUIREMETS .NET Core Development (hands-on core expertise) - C# / .NET Core Development Building / Maintaining Production-Grade Applications / APIs / Services Cloud Modernization - Migrating Legacy .NET Applications to Cloud-Native Azure Architectures Azure-Native Application Development - Designing / Deploying / Supporting .NET Applications Directly in Azure with Real Production Workloads (Beyond Portal-Based "Click-Ops") AKS / Containerization - Containerizing .NET Applications (Docker Multi-Stage Builds) Deploying / Managing within AKS (Deployments / Services / ConfigMaps / Secrets / Helm / Node Pools / Autoscaling / Ingress / Cluster-Level Troubleshooting) CI/CD / Azure DevOps - Building / Maintaining YAML-Based CI/CD Pipelines (Automated Builds / Testing / Security Scans / Docker Image Creation / Deployments to AKS / App Services) Azure Platform Services Integration - Integrating Core Azure Services into .NET Applications (App Services / Azure Functions / Service Bus / Event Grid / Azure Monitor / Application Insights With Custom Telemetry / Distributed Tracing / Metrics) IaC / Terraform - Developing / Managing Terraform-Based Infrastructure (AKS / VNETs / Private Endpoints / Service Bus / SQL Database / Storage) Aligned with Application Development Identity / Security (Azure) - Implementing Secure Patterns (Managed Identities / Entra ID / RBAC / Secure Connectivity / Polly Retry Policies / Private Endpoints / VNET Integration) within .NET Services Observability / SRE Practices - Applying SRE Principles within Development (Application Insights / Log Analytics / Alerting / Performance Monitoring / Production Troubleshooting) Database Design / Optimization - Oracle / MS SQL Server / NoSQL (CosmosDB) SQL Scripting (hands-on) Designing / Evolving Database Schemas Performing Query Performance Analysis Indexing to Deliver Scalable / Performant Services Problem Solving / Collaboration - Driving Root Cause Analysis / Debugging for .NET Applications in Azure Collaborating within SRE / Cross-Functional Teams PREFERRED (not required) Azure Governance / Best Practices - Azure Policy / GitHub Copilot Context / Embedding SRE Practices Early in Development Lifecycle Architecture Frameworks (Familiarity) - Azure Well-Architected Framework (Reliability / Operational Excellence) Advanced AKS / Hybrid - Exposure to Advanced AKS Topics (Cluster Autoscaling / Service Mesh / Cost Optimization) / Hybrid Scenarios (Azure Arc) AI / Automation Enablement - Driving Adoption of AI-Powered Tools (GitHub Copilot) to Improve Developer Productivity / Code Quality By submitting to this position, you are agreeing to be included in our talent pool for future hiring for similarly qualified positions. EEO Notice Vaco by Highspring is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of race (including but not limited to traits historically associated with race such as hair texture and hair style), color, sex (includes pregnancy or related conditions), religion or creed, national origin, citizenship, age, disability, status as a veteran, union membership, ethnicity, gender, gender identity, gender expression, sexual orientation, marital status, political affiliation, or any other protected characteristics as required by federal, state or local law. Vaco by Highspring and its parents, affiliates, and subsidiaries are committed to the full inclusion of all qualified individuals. As part of this commitment, Vaco by Highspring and its parents, affiliates, and subsidiaries will ensure that persons with disabilities are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact . Vaco by Highspring also wants all applicants to know their rights that workplace discrimination is illegal. Representation Notice By submitting to this position, you agree that you will be giving Vaco by Highspring the exclusive right to present your as a candidate for the foregoing employment opportunity. You further agree that you have represented information about yourself accurately and have not affirmatively misrepresented your qualifications. You also agree to maintain as confidential, to the fullest extent permitted by law, any information you learn from Vaco by Highspring about the position and you will limit disclosure of information about the position only to the extent necessary to perform any obligations in furtherance of your application. In exchange, Vaco by Highspring agrees to exercise reasonable efforts to represent you through all solicitation, job screening and resume dispersal. For residents of Ontario, Canada: Based on Highspring's discussions with its Client, Highspring's understanding is that this position for employment is a current vacancy (either through Highspring as a contractor or with the client directly). Privacy Notice Vaco by Highspring and its parents, affiliates, and subsidiaries ("we," "our," or "Vaco by Highspring") respects your privacy and are committed to providing transparent notice of our policies. California residents may access Vaco by Highspring HR Notice at Collection for California Applicants and Employees here. Virginia residents may access our state specific policies here. Residents of all other states may access our policies here. Canadian residents may access our policies in English here and in French here. Residents of countries governed by GDPR may access our policies here. Additionally, submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. More details about Vaco by Highspring's use of AI can be found here (). Further assessment of candidates beyond this initial phase will be conducted by recruiters and hiring managers. Vaco by Highspring does not know and cannot opine on if its client's use of AI products in hiring. Pay Transparency Notice Determining compensation for this role (and others) at Vaco by Highspring depends upon a wide array of factors including but not limited to: the individual's skill sets, experience and training; licensure and certification requirements; office location and other geographic considerations; other business and organizational needs. With that said, as required by local law, Vaco by Highspring believes that the following salary range referenced above reasonably estimates the base compensation for an individual hired into this position in geographies that require salary range disclosure. The individual may also be eligible for discretionary bonuses.
Vast is developing next-generation space stations and space infrastructure using an incremental, hardware-rich, and low-cost approach. Vast is rapidly developing its multi-module Haven Station to ensure a continuous human presence in space for America and its allies, enabling advanced microgravity research and manufacturing, and unlocking a new space economy for government, corporate, and private customers. Haven Demo's 2025 success made Vast the only operational commercial space station company to fly and operate its own spacecraft. Next, Haven-1 is expected to become the world's first commercial space station when it launches in 2027, followed by additional Haven modules. Additionally, the company recently announced Vast Satellite , a high-power satellite product line leveraging its space station components and the heritage of Haven Demo. Headquartered in Long Beach, California , and with more than 1,000 employees and over a billion dollars in private capital, Vast has built the facilities required to manufacture and operate America's next space station. The company plans to develop future habitats and systems for the Moon and Mars, dedicated space stations for government partners, and other crewed systems that will unlock the expanding long-term space economy. Vast is looking for a Staff Software Engineer, AI Tooling reporting to the Senior Manager, Software Engineering to support the development of the systems that will be required for the design and build of artificial-gravity human-rated space stations. This will be a full-time , exempt position located in our Long Beach location. Responsibilities: Define the technical direction and long-term roadmap for Vast's internal AI platforms and tooling Architect and lead the development of full-stack AI applications that serve diverse use-cases across the company Establish engineering best practices, design patterns, and quality standards for AI systems development Build and maintain production-grade integrations connecting AI models with internal tools, data sources, and workflows Design and implement pipelines for AI response enforcement, content safety, and output formatting Drive the evaluation and adoption of emerging AI frameworks, models, and patterns to continuously improve internal tooling Collaborate cross-functionally with IT Infrastructure, security, and business stakeholders to integrate AI tooling with existing systems Mentor and guide other engineers contributing to AI initiatives, fostering a culture of technical excellence Ensure the reliability, observability, and scalability of production AI systems Rapidly iterate on AI tooling as the technology landscape and company needs evolve Minimum Qualifications: Bachelor's degree in computer science, math, an engineering discipline, or equivalent work experience 8+ years of software development or relevant industry experience Proven track record of leading technical initiatives from concept through production deployment Strong full-stack development proficiency with experience across both frontend and backend systems Deep understanding of AI/ML concepts, large language model architectures, prompt engineering, and modern AI tooling patterns Demonstrated experience implementing and maintaining production applications at scale Experience establishing engineering best practices and providing technical mentorship Strong cross-functional communication skills with the ability to translate technical concepts for diverse audiences Preferred Skills & Experience: Experience with Python and TypeScript in production environments Experience with Svelte or SvelteKit Experience with OpenWebUI or similar open-source AI interface platforms Experience building or integrating MCP (Model Context Protocol) servers and gateways Experience with LLM API integration and AI pipeline development Experience with containerization and deployment (Docker, Kubernetes) Familiarity with RAG (Retrieval-Augmented Generation) patterns and vector databases Experience building or leading AI/ML platforms at scale Prior experience as a tech lead or in a similar engineering leadership role Experience working in fast-paced, ambiguous environments with evolving requirements Additional Requirements: Ability to travel up to 10% of the time Willingness to work evenings and/or weekends to support critical mission milestones Ability to lift up to 25 lbs unassisted Specific certifications, as appropriate Pay Range: Senior Software Engineer, AI Tooling: $159,900 - $226,900 Staff Software Engineer, AI Tooling: $188,600 - $267,700 Pay Range: California $159,900-$267,700 USD We calibrate level and compensation to the selected candidate's experience and demonstrated capability. We welcome applicants across a range of experience levels to apply. COMPENSATION AND BENEFITS Base salary will vary depending on job-related knowledge, education, skills, experience, business needs, and market demand. Salary is just one component of our comprehensive compensation package. Full-time employees also receive company equity, as well as access to a full suite of compelling benefits and perks, including: medical, dental, and vision coverage for employees and dependents, generous paid time off; up to 20+ days of vacation for exempt staff and up to 10+ days of vacation for non-exempt staff with the ability to cash-out unused vacation annually, paid parental leave, short and long-term disability insurance, life insurance, access to a 401(k) retirement plan, ClassPass credits, personalized mental healthcare through Spring Health, and other discounts and perks. We also take pride in offering exceptional food perks, with snacks, drip coffee & onsite barista, cold drinks, and dinner meals remaining free of charge, and lunch subsidized as part of Vast's ongoing commitment to providing high-quality meals for employees. U.S. EXPORT CONTROL COMPLIANCE STATUS The person hired will have access to information and items subject to U.S. export controls, and therefore, must either be a "U.S. person" as defined by 22 C.F.R. 120.62 or otherwise eligible for deemed export licensing. This status includes U.S. citizens, U.S. nationals, lawful permanent residents (green card holders), and asylees and refugees with such status granted, not pending. EQUAL OPPORTUNITY Vast is an Equal Opportunity Employer; employment with Vast is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.
09/25/2026
Full time
Vast is developing next-generation space stations and space infrastructure using an incremental, hardware-rich, and low-cost approach. Vast is rapidly developing its multi-module Haven Station to ensure a continuous human presence in space for America and its allies, enabling advanced microgravity research and manufacturing, and unlocking a new space economy for government, corporate, and private customers. Haven Demo's 2025 success made Vast the only operational commercial space station company to fly and operate its own spacecraft. Next, Haven-1 is expected to become the world's first commercial space station when it launches in 2027, followed by additional Haven modules. Additionally, the company recently announced Vast Satellite , a high-power satellite product line leveraging its space station components and the heritage of Haven Demo. Headquartered in Long Beach, California , and with more than 1,000 employees and over a billion dollars in private capital, Vast has built the facilities required to manufacture and operate America's next space station. The company plans to develop future habitats and systems for the Moon and Mars, dedicated space stations for government partners, and other crewed systems that will unlock the expanding long-term space economy. Vast is looking for a Staff Software Engineer, AI Tooling reporting to the Senior Manager, Software Engineering to support the development of the systems that will be required for the design and build of artificial-gravity human-rated space stations. This will be a full-time , exempt position located in our Long Beach location. Responsibilities: Define the technical direction and long-term roadmap for Vast's internal AI platforms and tooling Architect and lead the development of full-stack AI applications that serve diverse use-cases across the company Establish engineering best practices, design patterns, and quality standards for AI systems development Build and maintain production-grade integrations connecting AI models with internal tools, data sources, and workflows Design and implement pipelines for AI response enforcement, content safety, and output formatting Drive the evaluation and adoption of emerging AI frameworks, models, and patterns to continuously improve internal tooling Collaborate cross-functionally with IT Infrastructure, security, and business stakeholders to integrate AI tooling with existing systems Mentor and guide other engineers contributing to AI initiatives, fostering a culture of technical excellence Ensure the reliability, observability, and scalability of production AI systems Rapidly iterate on AI tooling as the technology landscape and company needs evolve Minimum Qualifications: Bachelor's degree in computer science, math, an engineering discipline, or equivalent work experience 8+ years of software development or relevant industry experience Proven track record of leading technical initiatives from concept through production deployment Strong full-stack development proficiency with experience across both frontend and backend systems Deep understanding of AI/ML concepts, large language model architectures, prompt engineering, and modern AI tooling patterns Demonstrated experience implementing and maintaining production applications at scale Experience establishing engineering best practices and providing technical mentorship Strong cross-functional communication skills with the ability to translate technical concepts for diverse audiences Preferred Skills & Experience: Experience with Python and TypeScript in production environments Experience with Svelte or SvelteKit Experience with OpenWebUI or similar open-source AI interface platforms Experience building or integrating MCP (Model Context Protocol) servers and gateways Experience with LLM API integration and AI pipeline development Experience with containerization and deployment (Docker, Kubernetes) Familiarity with RAG (Retrieval-Augmented Generation) patterns and vector databases Experience building or leading AI/ML platforms at scale Prior experience as a tech lead or in a similar engineering leadership role Experience working in fast-paced, ambiguous environments with evolving requirements Additional Requirements: Ability to travel up to 10% of the time Willingness to work evenings and/or weekends to support critical mission milestones Ability to lift up to 25 lbs unassisted Specific certifications, as appropriate Pay Range: Senior Software Engineer, AI Tooling: $159,900 - $226,900 Staff Software Engineer, AI Tooling: $188,600 - $267,700 Pay Range: California $159,900-$267,700 USD We calibrate level and compensation to the selected candidate's experience and demonstrated capability. We welcome applicants across a range of experience levels to apply. COMPENSATION AND BENEFITS Base salary will vary depending on job-related knowledge, education, skills, experience, business needs, and market demand. Salary is just one component of our comprehensive compensation package. Full-time employees also receive company equity, as well as access to a full suite of compelling benefits and perks, including: medical, dental, and vision coverage for employees and dependents, generous paid time off; up to 20+ days of vacation for exempt staff and up to 10+ days of vacation for non-exempt staff with the ability to cash-out unused vacation annually, paid parental leave, short and long-term disability insurance, life insurance, access to a 401(k) retirement plan, ClassPass credits, personalized mental healthcare through Spring Health, and other discounts and perks. We also take pride in offering exceptional food perks, with snacks, drip coffee & onsite barista, cold drinks, and dinner meals remaining free of charge, and lunch subsidized as part of Vast's ongoing commitment to providing high-quality meals for employees. U.S. EXPORT CONTROL COMPLIANCE STATUS The person hired will have access to information and items subject to U.S. export controls, and therefore, must either be a "U.S. person" as defined by 22 C.F.R. 120.62 or otherwise eligible for deemed export licensing. This status includes U.S. citizens, U.S. nationals, lawful permanent residents (green card holders), and asylees and refugees with such status granted, not pending. EQUAL OPPORTUNITY Vast is an Equal Opportunity Employer; employment with Vast is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.
Senior Staff AI Engineer (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership. Capital One is open to hiring a Remote Employee for this opportunity. Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Staff AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/24/2026
Full time
Senior Staff AI Engineer (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership. Capital One is open to hiring a Remote Employee for this opportunity. Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Staff AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Senior Staff AI Engineer (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership. Capital One is open to hiring a Remote Employee for this opportunity. Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Staff AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/24/2026
Full time
Senior Staff AI Engineer (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership. Capital One is open to hiring a Remote Employee for this opportunity. Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Remote (Regardless of Location): $286,200 - $326,700 for Sr. Staff AI Engineer Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
AI Engineer 4 (AI Foundations) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Sales Territory: $179,400 - $204,700 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/24/2026
Full time
AI Engineer 4 (AI Foundations) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Sales Territory: $179,400 - $204,700 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
AI Engineer 4 (AI Foundations) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Sales Territory: $179,400 - $204,700 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/24/2026
Full time
AI Engineer 4 (AI Foundations) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Sales Territory: $179,400 - $204,700 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Senior .NET Software Engineer / SRE DETAILS Location: Arlington, TX 76014 (hybrid onsite 2-days per week) Position Type: 6M C2H Hourly / Salary: to $160K+ (based on experience level) JOB SUMMARY Vaco is currently seeking a Senior .NET Software Engineer / SRE for a 6M C2H opportunity that is located in Arlington, TX 76014 (hybrid onsite 2-days per week). About the Project: This is an ongoing multi-year digital transformation initiative (now roughly 2+ years in) focused on modernizing financial services platforms, making them more scalable, reliable, automated, and cloud-native. The .NET Engineer / SRE will be building and extending tools / frameworks for SRE practices, including automation scripts, custom monitoring, CI/CD enhancements, reliability tooling, etc. The .NET Engineer / SRE will help bridge the gap between development teams (writing .NET) and SREs (ensuring reliability at scale), especially within this strategic / tool building-focused SRE practice. This is not a traditional .NET Developer (no owning user-facing APIs or business logic for financial services), but rather applying .NET skills to SRE tasks, including building custom tools for automation, reliability, and release engineering, creating GITHub Copilot context files to embed SRE practices (security / performance checks) early in development cycles, debugging / refactoring code in production environments, performing root-cause analysis, and iterating post-deployment. Additionally, there's no traditional on-call production monitoring / support, whereas the focus will be on design, architecture, automation, and enabling scalable DevOps / SRE practices for development teams, rather than day-to-day firefighting. Ideal Senior Candidate:The ideal candidate will come with a Developer-First Mindset for full SDLC, Architecture Ownership, and Problem Solving (build the tool if it does not exist approach). Having .NET fluency enables strategic SRE in their .NET / Azure ecosystem, bridging gaps between development and operations without siloed roles. The Senior Engineer MUST be able to walk through their most recent projects, in depth, describing complex features, debugging, and development frameworks (why you'd choose one over the other for particular tasks, etc.). JOB REQUIREMETS .NET Core Development (hands-on core expertise) - Strong C# / .NET (Core 6/8+) Development Building / Maintaining Production-Grade Applications / APIs / Services Clean / Maintainable Enterprise Code Design Patterns (Factory / Singleton, etc.) Onion / Clean Architecture Cloud Modernization / Migration - Migrating Legacy .NET Applications to Cloud-Native Azure Architectures Refactoring / Modernizing Existing Codebases While Maintaining High Uptime (99.9%+) Azure-Native Application Development - Designing / Deploying / Supporting .NET Applications Directly in Azure with Real Production Workloads Beyond Portal-Based "Click-Ops" Deep Integration of Azure Services into .NET Code AKS / Containerization - Containerizing .NET Applications (Docker Multi-Stage Builds) Deploying / Managing within AKS (Deployments / Services / ConfigMaps / Secrets / Helm / Node Pools/ Autoscaling / Ingress / Cluster-Level Troubleshooting) CI/CD / Release Engineering - Building / Maintaining YAML-Based CI/CD Pipelines in Azure DevOps Automated Builds / Testing / Security Scans / Docker Image Creation / Deployments to AKS / App Services Blue-Green Deployments / Feature Flags for Zero-Downtime Releases Azure Platform Services Integration - Integrating Core Azure Services into .NET Applications (App Services / Azure Functions / Service Bus / Event Grid / Azure Monitor / Application Insights with Custom Telemetry / Distributed Tracing / Metrics) IaC / Terraform - Developing / Managing Terraform-Based Infrastructure (AKS / VNETs / Private Endpoints / Service Bus / SQL Database / Storage) Aligned with Application Development Scripting (PowerShell / BASH / Python) Identity / Security (Azure) - Implementing Secure Patterns in .NET Services (Entra ID / Managed Identities / RBAC / OAuth 2.0 / JWT / Azure Key Vault / Polly Retry Policies / Private Endpoints / VNET Integration) SOX Compliance / Secure Authentication Flows Debugging Post-Migration Identity Issues (Claims / Tokens / Conditional Access) Observability / SRE Practices - Applying SRE Principles within Development Production Troubleshooting / Root-Cause Analysis / Iterative Improvements on .NET Applications in Azure Custom Telemetry / Logging / Performance Monitoring (Application Insights / Log Analytics / Alerting) Database Design / Optimization - Oracle / MS SQL Server / NoSQL (Cosmos DB) Hands-on SQL Scripting Designing / Evolving Database Schemas Query Performance Analysis / Indexing / Tuning (Post-Migration Behavior Changes, etc.) Problem Solving / Full SDLC Ownership - End-to-End Ownership of Complex Features / Improvements Architecture Decisions / Implementation / Production Debugging / Refactoring / Post-Deployment Enhancements Developer-First Mindset for Building Automation Tools / Reliability Frameworks PREFERRED (not required) Azure Governance / Best Practices - Azure Policy / GitHub Copilot Context / Embedding SRE Practices Early in Development Lifecycle Architecture Frameworks (Familiarity) - Azure Well-Architected Framework (Reliability / Operational Excellence) Advanced AKS / Hybrid - Exposure to Advanced AKS Topics (Cluster Autoscaling / Service Mesh / Cost Optimization) / Hybrid Scenarios (Azure Arc) AI / Automation Enablement - Driving Adoption of AI-Powered Tools (GitHub Copilot) to Improve Developer Productivity / Code Quality EEO Notice Vaco by Highspring is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of race (including but not limited to traits historically associated with race such as hair texture and hair style), color, sex (includes pregnancy or related conditions), religion or creed, national origin, citizenship, age, disability, status as a veteran, union membership, ethnicity, gender, gender identity, gender expression, sexual orientation, marital status, political affiliation, or any other protected characteristics as required by federal, state or local law. Vaco by Highspring and its parents, affiliates, and subsidiaries are committed to the full inclusion of all qualified individuals. As part of this commitment, Vaco by Highspring and its parents, affiliates, and subsidiaries will ensure that persons with disabilities are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact . Vaco by Highspring also wants all applicants to know their rights that workplace discrimination is illegal. Representation Notice By submitting to this position, you agree that you will be giving Vaco by Highspring the exclusive right to present your as a candidate for the foregoing employment opportunity. You further agree that you have represented information about yourself accurately and have not affirmatively misrepresented your qualifications. You also agree to maintain as confidential, to the fullest extent permitted by law, any information you learn from Vaco by Highspring about the position and you will limit disclosure of information about the position only to the extent necessary to perform any obligations in furtherance of your application. In exchange, Vaco by Highspring agrees to exercise reasonable efforts to represent you through all solicitation, job screening and resume dispersal. For residents of Ontario, Canada: Based on Highspring's discussions with its Client, Highspring's understanding is that this position for employment is a current vacancy (either through Highspring as a contractor or with the client directly). Privacy Notice Vaco by Highspring and its parents, affiliates, and subsidiaries ("we," "our," or "Vaco by Highspring") respects your privacy and are committed to providing transparent notice of our policies. California residents may access Vaco by Highspring HR Notice at Collection for California Applicants and Employees here. Virginia residents may access our state specific policies here. Residents of all other states may access our policies here. Canadian residents may access our policies in English here and in French here. Residents of countries governed by GDPR may access our policies here. Additionally, submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. More details about Vaco by Highspring's use of AI can be found here (). Further assessment of candidates beyond this initial phase will be conducted by recruiters and hiring managers. Vaco by Highspring does not know and cannot opine on if its client's use of AI products in hiring. Pay Transparency Notice Determining compensation for this role (and others) at Vaco by Highspring depends upon a wide array of factors including but not limited to: the individual's skill sets, experience and training; licensure and certification requirements; office location and other geographic considerations; other business and organizational needs. With that said, as required by local law, Vaco by Highspring believes that the following salary range referenced above reasonably estimates the base compensation for an individual hired into this position in geographies that require salary range disclosure. The individual may also be eligible for discretionary bonuses.
09/24/2026
Full time
Senior .NET Software Engineer / SRE DETAILS Location: Arlington, TX 76014 (hybrid onsite 2-days per week) Position Type: 6M C2H Hourly / Salary: to $160K+ (based on experience level) JOB SUMMARY Vaco is currently seeking a Senior .NET Software Engineer / SRE for a 6M C2H opportunity that is located in Arlington, TX 76014 (hybrid onsite 2-days per week). About the Project: This is an ongoing multi-year digital transformation initiative (now roughly 2+ years in) focused on modernizing financial services platforms, making them more scalable, reliable, automated, and cloud-native. The .NET Engineer / SRE will be building and extending tools / frameworks for SRE practices, including automation scripts, custom monitoring, CI/CD enhancements, reliability tooling, etc. The .NET Engineer / SRE will help bridge the gap between development teams (writing .NET) and SREs (ensuring reliability at scale), especially within this strategic / tool building-focused SRE practice. This is not a traditional .NET Developer (no owning user-facing APIs or business logic for financial services), but rather applying .NET skills to SRE tasks, including building custom tools for automation, reliability, and release engineering, creating GITHub Copilot context files to embed SRE practices (security / performance checks) early in development cycles, debugging / refactoring code in production environments, performing root-cause analysis, and iterating post-deployment. Additionally, there's no traditional on-call production monitoring / support, whereas the focus will be on design, architecture, automation, and enabling scalable DevOps / SRE practices for development teams, rather than day-to-day firefighting. Ideal Senior Candidate:The ideal candidate will come with a Developer-First Mindset for full SDLC, Architecture Ownership, and Problem Solving (build the tool if it does not exist approach). Having .NET fluency enables strategic SRE in their .NET / Azure ecosystem, bridging gaps between development and operations without siloed roles. The Senior Engineer MUST be able to walk through their most recent projects, in depth, describing complex features, debugging, and development frameworks (why you'd choose one over the other for particular tasks, etc.). JOB REQUIREMETS .NET Core Development (hands-on core expertise) - Strong C# / .NET (Core 6/8+) Development Building / Maintaining Production-Grade Applications / APIs / Services Clean / Maintainable Enterprise Code Design Patterns (Factory / Singleton, etc.) Onion / Clean Architecture Cloud Modernization / Migration - Migrating Legacy .NET Applications to Cloud-Native Azure Architectures Refactoring / Modernizing Existing Codebases While Maintaining High Uptime (99.9%+) Azure-Native Application Development - Designing / Deploying / Supporting .NET Applications Directly in Azure with Real Production Workloads Beyond Portal-Based "Click-Ops" Deep Integration of Azure Services into .NET Code AKS / Containerization - Containerizing .NET Applications (Docker Multi-Stage Builds) Deploying / Managing within AKS (Deployments / Services / ConfigMaps / Secrets / Helm / Node Pools/ Autoscaling / Ingress / Cluster-Level Troubleshooting) CI/CD / Release Engineering - Building / Maintaining YAML-Based CI/CD Pipelines in Azure DevOps Automated Builds / Testing / Security Scans / Docker Image Creation / Deployments to AKS / App Services Blue-Green Deployments / Feature Flags for Zero-Downtime Releases Azure Platform Services Integration - Integrating Core Azure Services into .NET Applications (App Services / Azure Functions / Service Bus / Event Grid / Azure Monitor / Application Insights with Custom Telemetry / Distributed Tracing / Metrics) IaC / Terraform - Developing / Managing Terraform-Based Infrastructure (AKS / VNETs / Private Endpoints / Service Bus / SQL Database / Storage) Aligned with Application Development Scripting (PowerShell / BASH / Python) Identity / Security (Azure) - Implementing Secure Patterns in .NET Services (Entra ID / Managed Identities / RBAC / OAuth 2.0 / JWT / Azure Key Vault / Polly Retry Policies / Private Endpoints / VNET Integration) SOX Compliance / Secure Authentication Flows Debugging Post-Migration Identity Issues (Claims / Tokens / Conditional Access) Observability / SRE Practices - Applying SRE Principles within Development Production Troubleshooting / Root-Cause Analysis / Iterative Improvements on .NET Applications in Azure Custom Telemetry / Logging / Performance Monitoring (Application Insights / Log Analytics / Alerting) Database Design / Optimization - Oracle / MS SQL Server / NoSQL (Cosmos DB) Hands-on SQL Scripting Designing / Evolving Database Schemas Query Performance Analysis / Indexing / Tuning (Post-Migration Behavior Changes, etc.) Problem Solving / Full SDLC Ownership - End-to-End Ownership of Complex Features / Improvements Architecture Decisions / Implementation / Production Debugging / Refactoring / Post-Deployment Enhancements Developer-First Mindset for Building Automation Tools / Reliability Frameworks PREFERRED (not required) Azure Governance / Best Practices - Azure Policy / GitHub Copilot Context / Embedding SRE Practices Early in Development Lifecycle Architecture Frameworks (Familiarity) - Azure Well-Architected Framework (Reliability / Operational Excellence) Advanced AKS / Hybrid - Exposure to Advanced AKS Topics (Cluster Autoscaling / Service Mesh / Cost Optimization) / Hybrid Scenarios (Azure Arc) AI / Automation Enablement - Driving Adoption of AI-Powered Tools (GitHub Copilot) to Improve Developer Productivity / Code Quality EEO Notice Vaco by Highspring is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of race (including but not limited to traits historically associated with race such as hair texture and hair style), color, sex (includes pregnancy or related conditions), religion or creed, national origin, citizenship, age, disability, status as a veteran, union membership, ethnicity, gender, gender identity, gender expression, sexual orientation, marital status, political affiliation, or any other protected characteristics as required by federal, state or local law. Vaco by Highspring and its parents, affiliates, and subsidiaries are committed to the full inclusion of all qualified individuals. As part of this commitment, Vaco by Highspring and its parents, affiliates, and subsidiaries will ensure that persons with disabilities are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact . Vaco by Highspring also wants all applicants to know their rights that workplace discrimination is illegal. Representation Notice By submitting to this position, you agree that you will be giving Vaco by Highspring the exclusive right to present your as a candidate for the foregoing employment opportunity. You further agree that you have represented information about yourself accurately and have not affirmatively misrepresented your qualifications. You also agree to maintain as confidential, to the fullest extent permitted by law, any information you learn from Vaco by Highspring about the position and you will limit disclosure of information about the position only to the extent necessary to perform any obligations in furtherance of your application. In exchange, Vaco by Highspring agrees to exercise reasonable efforts to represent you through all solicitation, job screening and resume dispersal. For residents of Ontario, Canada: Based on Highspring's discussions with its Client, Highspring's understanding is that this position for employment is a current vacancy (either through Highspring as a contractor or with the client directly). Privacy Notice Vaco by Highspring and its parents, affiliates, and subsidiaries ("we," "our," or "Vaco by Highspring") respects your privacy and are committed to providing transparent notice of our policies. California residents may access Vaco by Highspring HR Notice at Collection for California Applicants and Employees here. Virginia residents may access our state specific policies here. Residents of all other states may access our policies here. Canadian residents may access our policies in English here and in French here. Residents of countries governed by GDPR may access our policies here. Additionally, submissions to this position are subject to the use of AI to perform preliminary candidate screenings, focused on ensuring minimum job requirements noted in the position are satisfied. More details about Vaco by Highspring's use of AI can be found here (). Further assessment of candidates beyond this initial phase will be conducted by recruiters and hiring managers. Vaco by Highspring does not know and cannot opine on if its client's use of AI products in hiring. Pay Transparency Notice Determining compensation for this role (and others) at Vaco by Highspring depends upon a wide array of factors including but not limited to: the individual's skill sets, experience and training; licensure and certification requirements; office location and other geographic considerations; other business and organizational needs. With that said, as required by local law, Vaco by Highspring believes that the following salary range referenced above reasonably estimates the base compensation for an individual hired into this position in geographies that require salary range disclosure. The individual may also be eligible for discretionary bonuses.
Job DescriptionJob DescriptionJob Title: Artificial Intelligence Senior Associate Location: Dearborn, MI (local preferred) Duration: 12 Months (with potential for extension) Interview: Interview onsite in South Lyon/Novi, MI Candidate must be USC or GC. Do not apply for C2C. It will be a hybrid position. REVIEW JD MAKE SURE REQUIRED SKILLS (highlighted red) ARE ON THE RESUME. Will be onsite 4 days a week. Position Description: Employees in this job function are responsible for developing intelligent programs, cognitive applications and algorithms for data analysis and automation, leveraging various AI techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Key Responsibilities: Understand business requirements and develop AI algorithms, models and programs to solve complex problems, generate recommendations, extract patterns, make predictions, interpret sensor data (images, sound), orchestrate automation and enable self-service capabilities Perform large-scale experimentation and develop data driven applications that translate data into actionable intelligence Drive innovative applications of Artificial Intelligence tools and techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Research and optimize AI technologies to enhance efficiency and accuracy of data analysis and create more efficient automation Skills Required: Google Cloud Platform Experience Required: Bachelor's or Master's degree in Computer Science, Software Engineering, or related field (or equivalent practical experience). 3+ years building production software systems, including 1-2+ years on ML/AI or LLM-based applications. Proven experience designing and deploying multi-agent or multi-service architectures in production - not just notebooks or demos. As one 2026 hiring analysis puts it, the job is closer to distributed systems engineering with a probabilistic component than it is to ML research or prompt tweaking . Strong Python proficiency, including async/concurrent programming, and experience with backend frameworks (FastAPI, Flask). Hands-on experience with agent orchestration frameworks - LangGraph, CrewAI, LlamaIndex, or equivalent - for building stateful, multi-step, tool-using agent workflows. Practical experience building RAG pipelines: vector databases (pgvector, Pinecone, Weaviate, or Qdrant), embeddings, chunking strategies, and retrieval evaluation. Cloud deployment experience, ideally Google Cloud Platform (BigQuery, Cloud Run/GKE, Vertex AI, Pub/Sub) or equivalent AWS/Azure services. Strong SQL skills and experience with cloud data warehouses. Containerization and CI/CD experience (Docker, Kubernetes, GitHub Actions/Cloud Build). Experience building evaluation and observability pipelines for LLM/agent systems - offline eval sets, LLM-as-judge scoring, and tracing tools (LangSmith, Langfuse, OpenTelemetry, or equivalent) to track task success, latency, and cost. Understanding of LLM safety practices: guardrails, output validation, prompt-injection defense, and safe execution of AI-generated code/SQL (sandboxing, least privilege). Solid software engineering fundamentals: API design, testing, version control, security best practices. Experience Preferred: Experience with cost optimization and model routing - designing tiered pipelines that route between low-cost and high-capability models based on task complexity, and modeling per-conversation or per-task cost at scale. Experience deploying agentic systems with human-in-the-loop or multi-checkpoint validation workflows for high-reliability/high-stakes use cases. Experience with automotive, EV charging, IoT, or connected-vehicle telemetry data. Familiarity with Model Context Protocol (MCP) or similar standards for tool/data integration across agents. Prior experience in a startup or 0-to-1 product environment, comfortable with ambiguity and fast-evolving requirements. Education Required: Bachelor's Degree Education Preferred: Master's Degree Additional Information: Architect and deploy the production multi-agent orchestration layer (interpreter/orchestrator, NL-to-SQL agent, visualization agent, RCA/RAG agent, report composition agent, notification agent), using modern agent frameworks with state management and checkpointing rather than ad-hoc loops. Design and productionize RAG pipelines (chunking, embeddings, hybrid retrieval, reranking) grounded in approved schemas, engineering documentation, and historical issue records. Own BigQuery integration and enforce safe, least-privilege, validated execution of LLM-generated SQL. Build CI/CD, containerization, and infrastructure-as-code for deploying agent services on GCP (Cloud Run/GKE, Vertex AI). Implement evaluation pipelines and observability/tracing for every agent (golden datasets, LLM-as-judge scoring, regression alerts) so quality is measurable, not assumed. Implement guardrails, prompt-injection defenses, and human-in-the-loop approval checkpoints to ensure correctness and safety before any output triggers downstream action. Design cost/latency optimization strategies, including tiered model routing (cheap filter models vs. high-capability deep-dive models) and caching. Integrate validated outputs with operational systems (Salesforce ticketing, driver/site-manager notifications) and report export pipelines (PDF/HTML/spreadsheet). Collaborate with data scientists to productionize prototypes (anomaly detection, diagnostic agents) into scalable, monitored services. Establish versioning, testing, and safe rollout practices (canary/shadow deployments) for evolving agent logic.
09/23/2026
Full time
Job DescriptionJob DescriptionJob Title: Artificial Intelligence Senior Associate Location: Dearborn, MI (local preferred) Duration: 12 Months (with potential for extension) Interview: Interview onsite in South Lyon/Novi, MI Candidate must be USC or GC. Do not apply for C2C. It will be a hybrid position. REVIEW JD MAKE SURE REQUIRED SKILLS (highlighted red) ARE ON THE RESUME. Will be onsite 4 days a week. Position Description: Employees in this job function are responsible for developing intelligent programs, cognitive applications and algorithms for data analysis and automation, leveraging various AI techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Key Responsibilities: Understand business requirements and develop AI algorithms, models and programs to solve complex problems, generate recommendations, extract patterns, make predictions, interpret sensor data (images, sound), orchestrate automation and enable self-service capabilities Perform large-scale experimentation and develop data driven applications that translate data into actionable intelligence Drive innovative applications of Artificial Intelligence tools and techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Research and optimize AI technologies to enhance efficiency and accuracy of data analysis and create more efficient automation Skills Required: Google Cloud Platform Experience Required: Bachelor's or Master's degree in Computer Science, Software Engineering, or related field (or equivalent practical experience). 3+ years building production software systems, including 1-2+ years on ML/AI or LLM-based applications. Proven experience designing and deploying multi-agent or multi-service architectures in production - not just notebooks or demos. As one 2026 hiring analysis puts it, the job is closer to distributed systems engineering with a probabilistic component than it is to ML research or prompt tweaking . Strong Python proficiency, including async/concurrent programming, and experience with backend frameworks (FastAPI, Flask). Hands-on experience with agent orchestration frameworks - LangGraph, CrewAI, LlamaIndex, or equivalent - for building stateful, multi-step, tool-using agent workflows. Practical experience building RAG pipelines: vector databases (pgvector, Pinecone, Weaviate, or Qdrant), embeddings, chunking strategies, and retrieval evaluation. Cloud deployment experience, ideally Google Cloud Platform (BigQuery, Cloud Run/GKE, Vertex AI, Pub/Sub) or equivalent AWS/Azure services. Strong SQL skills and experience with cloud data warehouses. Containerization and CI/CD experience (Docker, Kubernetes, GitHub Actions/Cloud Build). Experience building evaluation and observability pipelines for LLM/agent systems - offline eval sets, LLM-as-judge scoring, and tracing tools (LangSmith, Langfuse, OpenTelemetry, or equivalent) to track task success, latency, and cost. Understanding of LLM safety practices: guardrails, output validation, prompt-injection defense, and safe execution of AI-generated code/SQL (sandboxing, least privilege). Solid software engineering fundamentals: API design, testing, version control, security best practices. Experience Preferred: Experience with cost optimization and model routing - designing tiered pipelines that route between low-cost and high-capability models based on task complexity, and modeling per-conversation or per-task cost at scale. Experience deploying agentic systems with human-in-the-loop or multi-checkpoint validation workflows for high-reliability/high-stakes use cases. Experience with automotive, EV charging, IoT, or connected-vehicle telemetry data. Familiarity with Model Context Protocol (MCP) or similar standards for tool/data integration across agents. Prior experience in a startup or 0-to-1 product environment, comfortable with ambiguity and fast-evolving requirements. Education Required: Bachelor's Degree Education Preferred: Master's Degree Additional Information: Architect and deploy the production multi-agent orchestration layer (interpreter/orchestrator, NL-to-SQL agent, visualization agent, RCA/RAG agent, report composition agent, notification agent), using modern agent frameworks with state management and checkpointing rather than ad-hoc loops. Design and productionize RAG pipelines (chunking, embeddings, hybrid retrieval, reranking) grounded in approved schemas, engineering documentation, and historical issue records. Own BigQuery integration and enforce safe, least-privilege, validated execution of LLM-generated SQL. Build CI/CD, containerization, and infrastructure-as-code for deploying agent services on GCP (Cloud Run/GKE, Vertex AI). Implement evaluation pipelines and observability/tracing for every agent (golden datasets, LLM-as-judge scoring, regression alerts) so quality is measurable, not assumed. Implement guardrails, prompt-injection defenses, and human-in-the-loop approval checkpoints to ensure correctness and safety before any output triggers downstream action. Design cost/latency optimization strategies, including tiered model routing (cheap filter models vs. high-capability deep-dive models) and caching. Integrate validated outputs with operational systems (Salesforce ticketing, driver/site-manager notifications) and report export pipelines (PDF/HTML/spreadsheet). Collaborate with data scientists to productionize prototypes (anomaly detection, diagnostic agents) into scalable, monitored services. Establish versioning, testing, and safe rollout practices (canary/shadow deployments) for evolving agent logic.
Job DescriptionJob DescriptionJob Title: Artificial Intelligence Senior Associate Location: Dearborn, MI (local preferred) Duration: 12 Months (with potential for extension) Interview: Interview onsite in South Lyon/Novi, MI Candidate must be USC or GC. Do not apply for C2C. It will be a hybrid position. REVIEW JD MAKE SURE REQUIRED SKILLS (highlighted red) ARE ON THE RESUME. Will be onsite 4 days a week. Position Description: Employees in this job function are responsible for developing intelligent programs, cognitive applications and algorithms for data analysis and automation, leveraging various AI techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Key Responsibilities: Understand business requirements and develop AI algorithms, models and programs to solve complex problems, generate recommendations, extract patterns, make predictions, interpret sensor data (images, sound), orchestrate automation and enable self-service capabilities Perform large-scale experimentation and develop data driven applications that translate data into actionable intelligence Drive innovative applications of Artificial Intelligence tools and techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Research and optimize AI technologies to enhance efficiency and accuracy of data analysis and create more efficient automation Skills Required: Google Cloud Platform Experience Required: Bachelor's or Master's degree in Computer Science, Software Engineering, or related field (or equivalent practical experience). 3+ years building production software systems, including 1-2+ years on ML/AI or LLM-based applications. Proven experience designing and deploying multi-agent or multi-service architectures in production - not just notebooks or demos. As one 2026 hiring analysis puts it, the job is closer to distributed systems engineering with a probabilistic component than it is to ML research or prompt tweaking . Strong Python proficiency, including async/concurrent programming, and experience with backend frameworks (FastAPI, Flask). Hands-on experience with agent orchestration frameworks - LangGraph, CrewAI, LlamaIndex, or equivalent - for building stateful, multi-step, tool-using agent workflows. Practical experience building RAG pipelines: vector databases (pgvector, Pinecone, Weaviate, or Qdrant), embeddings, chunking strategies, and retrieval evaluation. Cloud deployment experience, ideally Google Cloud Platform (BigQuery, Cloud Run/GKE, Vertex AI, Pub/Sub) or equivalent AWS/Azure services. Strong SQL skills and experience with cloud data warehouses. Containerization and CI/CD experience (Docker, Kubernetes, GitHub Actions/Cloud Build). Experience building evaluation and observability pipelines for LLM/agent systems - offline eval sets, LLM-as-judge scoring, and tracing tools (LangSmith, Langfuse, OpenTelemetry, or equivalent) to track task success, latency, and cost. Understanding of LLM safety practices: guardrails, output validation, prompt-injection defense, and safe execution of AI-generated code/SQL (sandboxing, least privilege). Solid software engineering fundamentals: API design, testing, version control, security best practices. Experience Preferred: Experience with cost optimization and model routing - designing tiered pipelines that route between low-cost and high-capability models based on task complexity, and modeling per-conversation or per-task cost at scale. Experience deploying agentic systems with human-in-the-loop or multi-checkpoint validation workflows for high-reliability/high-stakes use cases. Experience with automotive, EV charging, IoT, or connected-vehicle telemetry data. Familiarity with Model Context Protocol (MCP) or similar standards for tool/data integration across agents. Prior experience in a startup or 0-to-1 product environment, comfortable with ambiguity and fast-evolving requirements. Education Required: Bachelor's Degree Education Preferred: Master's Degree Additional Information: Architect and deploy the production multi-agent orchestration layer (interpreter/orchestrator, NL-to-SQL agent, visualization agent, RCA/RAG agent, report composition agent, notification agent), using modern agent frameworks with state management and checkpointing rather than ad-hoc loops. Design and productionize RAG pipelines (chunking, embeddings, hybrid retrieval, reranking) grounded in approved schemas, engineering documentation, and historical issue records. Own BigQuery integration and enforce safe, least-privilege, validated execution of LLM-generated SQL. Build CI/CD, containerization, and infrastructure-as-code for deploying agent services on GCP (Cloud Run/GKE, Vertex AI). Implement evaluation pipelines and observability/tracing for every agent (golden datasets, LLM-as-judge scoring, regression alerts) so quality is measurable, not assumed. Implement guardrails, prompt-injection defenses, and human-in-the-loop approval checkpoints to ensure correctness and safety before any output triggers downstream action. Design cost/latency optimization strategies, including tiered model routing (cheap filter models vs. high-capability deep-dive models) and caching. Integrate validated outputs with operational systems (Salesforce ticketing, driver/site-manager notifications) and report export pipelines (PDF/HTML/spreadsheet). Collaborate with data scientists to productionize prototypes (anomaly detection, diagnostic agents) into scalable, monitored services. Establish versioning, testing, and safe rollout practices (canary/shadow deployments) for evolving agent logic.
09/23/2026
Full time
Job DescriptionJob DescriptionJob Title: Artificial Intelligence Senior Associate Location: Dearborn, MI (local preferred) Duration: 12 Months (with potential for extension) Interview: Interview onsite in South Lyon/Novi, MI Candidate must be USC or GC. Do not apply for C2C. It will be a hybrid position. REVIEW JD MAKE SURE REQUIRED SKILLS (highlighted red) ARE ON THE RESUME. Will be onsite 4 days a week. Position Description: Employees in this job function are responsible for developing intelligent programs, cognitive applications and algorithms for data analysis and automation, leveraging various AI techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Key Responsibilities: Understand business requirements and develop AI algorithms, models and programs to solve complex problems, generate recommendations, extract patterns, make predictions, interpret sensor data (images, sound), orchestrate automation and enable self-service capabilities Perform large-scale experimentation and develop data driven applications that translate data into actionable intelligence Drive innovative applications of Artificial Intelligence tools and techniques such as deep learning, generative AI, natural language processing, image processing, cognitive automation, intelligent process automation, reinforcement learning, virtual assistants and specialized programming Research and optimize AI technologies to enhance efficiency and accuracy of data analysis and create more efficient automation Skills Required: Google Cloud Platform Experience Required: Bachelor's or Master's degree in Computer Science, Software Engineering, or related field (or equivalent practical experience). 3+ years building production software systems, including 1-2+ years on ML/AI or LLM-based applications. Proven experience designing and deploying multi-agent or multi-service architectures in production - not just notebooks or demos. As one 2026 hiring analysis puts it, the job is closer to distributed systems engineering with a probabilistic component than it is to ML research or prompt tweaking . Strong Python proficiency, including async/concurrent programming, and experience with backend frameworks (FastAPI, Flask). Hands-on experience with agent orchestration frameworks - LangGraph, CrewAI, LlamaIndex, or equivalent - for building stateful, multi-step, tool-using agent workflows. Practical experience building RAG pipelines: vector databases (pgvector, Pinecone, Weaviate, or Qdrant), embeddings, chunking strategies, and retrieval evaluation. Cloud deployment experience, ideally Google Cloud Platform (BigQuery, Cloud Run/GKE, Vertex AI, Pub/Sub) or equivalent AWS/Azure services. Strong SQL skills and experience with cloud data warehouses. Containerization and CI/CD experience (Docker, Kubernetes, GitHub Actions/Cloud Build). Experience building evaluation and observability pipelines for LLM/agent systems - offline eval sets, LLM-as-judge scoring, and tracing tools (LangSmith, Langfuse, OpenTelemetry, or equivalent) to track task success, latency, and cost. Understanding of LLM safety practices: guardrails, output validation, prompt-injection defense, and safe execution of AI-generated code/SQL (sandboxing, least privilege). Solid software engineering fundamentals: API design, testing, version control, security best practices. Experience Preferred: Experience with cost optimization and model routing - designing tiered pipelines that route between low-cost and high-capability models based on task complexity, and modeling per-conversation or per-task cost at scale. Experience deploying agentic systems with human-in-the-loop or multi-checkpoint validation workflows for high-reliability/high-stakes use cases. Experience with automotive, EV charging, IoT, or connected-vehicle telemetry data. Familiarity with Model Context Protocol (MCP) or similar standards for tool/data integration across agents. Prior experience in a startup or 0-to-1 product environment, comfortable with ambiguity and fast-evolving requirements. Education Required: Bachelor's Degree Education Preferred: Master's Degree Additional Information: Architect and deploy the production multi-agent orchestration layer (interpreter/orchestrator, NL-to-SQL agent, visualization agent, RCA/RAG agent, report composition agent, notification agent), using modern agent frameworks with state management and checkpointing rather than ad-hoc loops. Design and productionize RAG pipelines (chunking, embeddings, hybrid retrieval, reranking) grounded in approved schemas, engineering documentation, and historical issue records. Own BigQuery integration and enforce safe, least-privilege, validated execution of LLM-generated SQL. Build CI/CD, containerization, and infrastructure-as-code for deploying agent services on GCP (Cloud Run/GKE, Vertex AI). Implement evaluation pipelines and observability/tracing for every agent (golden datasets, LLM-as-judge scoring, regression alerts) so quality is measurable, not assumed. Implement guardrails, prompt-injection defenses, and human-in-the-loop approval checkpoints to ensure correctness and safety before any output triggers downstream action. Design cost/latency optimization strategies, including tiered model routing (cheap filter models vs. high-capability deep-dive models) and caching. Integrate validated outputs with operational systems (Salesforce ticketing, driver/site-manager notifications) and report export pipelines (PDF/HTML/spreadsheet). Collaborate with data scientists to productionize prototypes (anomaly detection, diagnostic agents) into scalable, monitored services. Establish versioning, testing, and safe rollout practices (canary/shadow deployments) for evolving agent logic.
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
AI Engineer 4 (AI Foundations, LLM Customization and Finetuning) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Francisco, CA: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
AI Engineer 4 (AI Foundations, LLM Customization and Finetuning) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Francisco, CA: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
Sr. Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Define and steer the technical AI architecture vision, integrating applied research breakthroughs into production ecosystems with reliability and scale Lead the establishment of AI performance, safety, and transparency standards that guide all model development and deployment company-wide Drive multi-year platform initiatives that unify data, compute and model lifecycle management under and cohesive enterprise AI architecture Mentor senior technical leaders across research, data and engineering disciplines, developing the next generation of Capital One's AI technical leadership Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 10 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies At least 10 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience architecting AI platforms with tradeoff decisions around cost, latency, throughput and accuracy 9 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience architecting, designing, developing, integrating, delivering, and supporting complex AI systems Demonstrated ability to lead and mentor multiple engineering teams and influence cross-functional stakeholders up to the SVP level Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Excellent communication and presentation skills, with the ability to articulate complex AI concepts to peers Recognized industry leader in applied AI or machine learning infrastructure through patents, publications, or open-source leadership Demonstrated experience designing long-term AI infrastructure strategies - balancing cost, scale, ethics and regulatory compliance Experience driving organization-wide adoption of AI safety, alignment and governance standards, collaborating with policy, risk and legal teams Proven ability to shape R&D investment strategy by identifying breakthrough AI capabilities with material business impact Experience right-sizing models, instance counts, and hardware types given requirements (e.g., context length, token inputs, token outputs) Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $314,800 - $359,300 for Sr. Staff AI Engineer McLean, VA: $314,800 - $359,300 for Sr. Staff AI Engineer New York, NY: $343,400 - $392,000 for Sr. Staff AI Engineer Richmond, VA: $286,200 - $326,700 for Sr. Staff AI Engineer San Francisco, CA: $343,400 - $392,000 for Sr. Staff AI Engineer San Jose, CA: $343,400 - $392,000 for Sr. Staff AI Engineer Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
AI Engineer 4 (AI Foundations, LLM Customization and Finetuning) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Francisco, CA: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
09/23/2026
Full time
AI Engineer 4 (AI Foundations, LLM Customization and Finetuning) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine learning - position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Intelligent Foundations and Experiences (IFX) team is at the center of bringing our vision for AI at Capital One to life. We work hand-in-hand with our partners across the company to advance the state of the art in science and AI engineering, and we build and deploy proprietary solutions that are central to our business and deliver value to millions of customers. Our AI models and platforms empower teams across Capital One to enhance their products with the transformative power of AI, in responsible and scalable ways for the highest leverage impact. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art foundation model optimization techniques to improve the performance - scalability, cost, latency, throughput - of large scale production AI systems. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize GPU/TPU utilization and accelerate model inference pipelines Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met Mentor Principal and Senior Associates on scalable design, performance tuning and research-to-production translation The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Basic Qualifications: Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies At least 4 years of experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications: Experience leading development AI systems with tradeoff decisions around cost, latency, throughput and accuracy 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g. AWS, Google Cloud, Azure, or equivalent private cloud) Experience designing, developing, delivering, and supporting AI services Experience developing AI and ML algorithms or technologies (e.g. LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, CUDA, or Golang Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost Experience in building agentic AI systems and agentic workflows Passion for staying abreast of the latest AI research and AI systems, and judiciously apply novel techniques in production Proficiency in designing distributed systems for model training, evaluation, and online inference at petabyte scale Experience defining AI model governance processes, including producibility, lineage tracking, and automated retaining schedules Demonstrated ability to influence architectural decisions across multiple AI product lines or platforms Capital One will consider sponsoring a new qualified applicant for employment authorization for this position. The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked. Cambridge, MA: $197,300 - $225,100 for AI Engineer 4 McLean, VA: $197,300 - $225,100 for AI Engineer 4 New York, NY: $215,200 - $245,600 for AI Engineer 4 San Francisco, CA: $215,200 - $245,600 for AI Engineer 4 San Jose, CA: $215,200 - $245,600 for AI Engineer 4 Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate's offer letter. This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan. Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website . Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level. This role is expected to accept applications for a minimum of 5 business days.No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history in a manner consistent with the requirements of applicable laws regarding criminal background inquiries, including, to the extent applicable, Article 23-A of the New York Correction Law; San Francisco, California Police Code Article 49, Sections ; New York City's Fair Chance Act; Philadelphia's Fair Criminal Records Screening Act; and other applicable federal, state, and local laws and regulations regarding criminal background inquiries. If you have visited our website in search of information on employment opportunities or to apply for a position, and you require an accommodation, please contact Capital One Recruiting at 1- or via email at . All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations. For technical support or questions about Capital One's recruiting process, please send an email to Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe and any position posted in the Philippines is for Capital One Philippines Service Corp. (COPSSC).
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an experienced Senior Operations Engineer to join the MetalDev Operations team. Reporting to the Engineering Manager, this hands-on technical role focuses on service reliability, observability, and operational excellence. You will develop deep expertise in the Redfish-based services and automation tools used by frontline operations teams during data center bring-ups and in production environments. You will identify gaps in tooling and operational processes, partner with the engineering team to develop and validate fixes, and help ensure our automation operates reliably at scale. You will also submit change requests to hardware and firmware vendors, validate vendor-provided fixes, and continuously improve runbooks and on-call documentation to reduce recurring incidents and operational toil. In this role, you will work with cutting-edge infrastructure, including high-performance NVIDIA GPU servers, Cooling Distribution Units (CDUs), NVLink switches, and power shelves supporting GB200 and GB300 Vera Rubin NVL72 systems and custom in-house hardware. You will collaborate closely with the Hardware Engineering, Fleet Operations, and the Service Engineering teams. Approximately 80% of the role will focus on day-to-day operations, production support, and incident response. The remaining 20% will focus on improving operational processes, observability, documentation, and remediation capabilities and automation to prevent recurring incidents and increase service reliability. Key Responsibilities Triage and Troubleshooting Troubleshoot the team owned services, including initialization and reboot issues. Diagnose problems involving BMCs of, servers, DPUs, power shelves, Cooling Distribution Units (CDUs). Monitor fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Partner with Fleet Operations and other engineering teams to resolve complex or recurring production issues. Perform root-cause analysis and post-incident reviews, ensuring corrective actions are documented, tracked, and completed. Maintain clear incident communications, runbooks, escalation procedures, and operational records. Support the team owned services in production environments and participate in the team's on-call rotation. Observability, Reliability, and Vendor Partnerships Own monitoring, dashboards, alerts, and operational KPIs for the team owned services using Prometheus and Grafana. Define reliability objectives and drive measurable reductions in incidents, escalations, and recurring support requests. Investigate hardware, firmware, and issues in partnership with internal engineering teams and external vendors. Manage vendor support cases involving BMCs, servers, power systems, and cooling infrastructure. Collect and provide diagnostic data, track issues through resolution, and validate vendor fixes before production rollout. Documentation and Continuous Improvement Create and maintain operational documentation, troubleshooting guides, escalation procedures, and service-support materials. Capture and share incident findings and operational knowledge across the team and partner teams. Identify repetitive operational tasks and develop more efficient, consistent remediation processes. Evaluate team processes using operational data and incident trends, and recommend improvements. Minimum Qualifications 5+ of experience in cloud operations, site reliability engineering (SRE), infrastructure operations, or a related technical field. Working knowledge of Kubernetes and at least one public cloud platform, such as AWS or GCP. Experience deploying and supporting containerized applications in Kubernetes environments. Experience with incident management practices, including incident response, escalation, and post-incident review processes. Experience using Prometheus, Grafana, and PromQL for monitoring, alerting, and troubleshooting. Strong knowledge of Linux system administration and internals and scripting. Experience troubleshooting complex issues across software services, operating systems, networks, and physical infrastructure. Experience participating in an on-call rotation supporting production services. Strong analytical and problem-solving skills, with a methodical approach to troubleshooting. Excellent written and verbal communication skills, particularly during high-impact incidents. Strong documentation skills and attention to detail. Preferred Qualifications Experience with server hardware, BMCs, Redfish, IPMI, or hardware-management services. Experience troubleshooting server provisioning, reboot, provisioning, or lifecycle-management failures. Familiarity with high-performance computing, GPU infrastructure, DPUs, or large-scale AI clusters. Experience working in data center environments, including server racks, power-distribution equipment, and cooling systems. Experience collaborating directly with hardware or firmware vendors to qualify and validate fixes. Bachelor's degree in computer science, engineering, or a related discipline-or equivalent practical experience. Understanding of Python or Golang. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You enjoy working close to the hardware and are curious about how GPUs, servers, and data centers fit together. You thrive in infrastructure environments where reliability, performance, and automation matter as much as features. You like collaborating across hardware, platform, and product teams to solve complex, ambiguous problems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $134,000 to $179,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace . click apply for full job details
09/23/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an experienced Senior Operations Engineer to join the MetalDev Operations team. Reporting to the Engineering Manager, this hands-on technical role focuses on service reliability, observability, and operational excellence. You will develop deep expertise in the Redfish-based services and automation tools used by frontline operations teams during data center bring-ups and in production environments. You will identify gaps in tooling and operational processes, partner with the engineering team to develop and validate fixes, and help ensure our automation operates reliably at scale. You will also submit change requests to hardware and firmware vendors, validate vendor-provided fixes, and continuously improve runbooks and on-call documentation to reduce recurring incidents and operational toil. In this role, you will work with cutting-edge infrastructure, including high-performance NVIDIA GPU servers, Cooling Distribution Units (CDUs), NVLink switches, and power shelves supporting GB200 and GB300 Vera Rubin NVL72 systems and custom in-house hardware. You will collaborate closely with the Hardware Engineering, Fleet Operations, and the Service Engineering teams. Approximately 80% of the role will focus on day-to-day operations, production support, and incident response. The remaining 20% will focus on improving operational processes, observability, documentation, and remediation capabilities and automation to prevent recurring incidents and increase service reliability. Key Responsibilities Triage and Troubleshooting Troubleshoot the team owned services, including initialization and reboot issues. Diagnose problems involving BMCs of, servers, DPUs, power shelves, Cooling Distribution Units (CDUs). Monitor fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Partner with Fleet Operations and other engineering teams to resolve complex or recurring production issues. Perform root-cause analysis and post-incident reviews, ensuring corrective actions are documented, tracked, and completed. Maintain clear incident communications, runbooks, escalation procedures, and operational records. Support the team owned services in production environments and participate in the team's on-call rotation. Observability, Reliability, and Vendor Partnerships Own monitoring, dashboards, alerts, and operational KPIs for the team owned services using Prometheus and Grafana. Define reliability objectives and drive measurable reductions in incidents, escalations, and recurring support requests. Investigate hardware, firmware, and issues in partnership with internal engineering teams and external vendors. Manage vendor support cases involving BMCs, servers, power systems, and cooling infrastructure. Collect and provide diagnostic data, track issues through resolution, and validate vendor fixes before production rollout. Documentation and Continuous Improvement Create and maintain operational documentation, troubleshooting guides, escalation procedures, and service-support materials. Capture and share incident findings and operational knowledge across the team and partner teams. Identify repetitive operational tasks and develop more efficient, consistent remediation processes. Evaluate team processes using operational data and incident trends, and recommend improvements. Minimum Qualifications 5+ of experience in cloud operations, site reliability engineering (SRE), infrastructure operations, or a related technical field. Working knowledge of Kubernetes and at least one public cloud platform, such as AWS or GCP. Experience deploying and supporting containerized applications in Kubernetes environments. Experience with incident management practices, including incident response, escalation, and post-incident review processes. Experience using Prometheus, Grafana, and PromQL for monitoring, alerting, and troubleshooting. Strong knowledge of Linux system administration and internals and scripting. Experience troubleshooting complex issues across software services, operating systems, networks, and physical infrastructure. Experience participating in an on-call rotation supporting production services. Strong analytical and problem-solving skills, with a methodical approach to troubleshooting. Excellent written and verbal communication skills, particularly during high-impact incidents. Strong documentation skills and attention to detail. Preferred Qualifications Experience with server hardware, BMCs, Redfish, IPMI, or hardware-management services. Experience troubleshooting server provisioning, reboot, provisioning, or lifecycle-management failures. Familiarity with high-performance computing, GPU infrastructure, DPUs, or large-scale AI clusters. Experience working in data center environments, including server racks, power-distribution equipment, and cooling systems. Experience collaborating directly with hardware or firmware vendors to qualify and validate fixes. Bachelor's degree in computer science, engineering, or a related discipline-or equivalent practical experience. Understanding of Python or Golang. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You enjoy working close to the hardware and are curious about how GPUs, servers, and data centers fit together. You thrive in infrastructure environments where reliability, performance, and automation matter as much as features. You like collaborating across hardware, platform, and product teams to solve complex, ambiguous problems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $134,000 to $179,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace . click apply for full job details
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an experienced Senior Operations Engineer to join the MetalDev Operations team. Reporting to the Engineering Manager, this hands-on technical role focuses on service reliability, observability, and operational excellence. You will develop deep expertise in the Redfish-based services and automation tools used by frontline operations teams during data center bring-ups and in production environments. You will identify gaps in tooling and operational processes, partner with the engineering team to develop and validate fixes, and help ensure our automation operates reliably at scale. You will also submit change requests to hardware and firmware vendors, validate vendor-provided fixes, and continuously improve runbooks and on-call documentation to reduce recurring incidents and operational toil. In this role, you will work with cutting-edge infrastructure, including high-performance NVIDIA GPU servers, Cooling Distribution Units (CDUs), NVLink switches, and power shelves supporting GB200 and GB300 Vera Rubin NVL72 systems and custom in-house hardware. You will collaborate closely with the Hardware Engineering, Fleet Operations, and the Service Engineering teams. Approximately 80% of the role will focus on day-to-day operations, production support, and incident response. The remaining 20% will focus on improving operational processes, observability, documentation, and remediation capabilities and automation to prevent recurring incidents and increase service reliability. Key Responsibilities Triage and Troubleshooting Troubleshoot the team owned services, including initialization and reboot issues. Diagnose problems involving BMCs of, servers, DPUs, power shelves, Cooling Distribution Units (CDUs). Monitor fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Partner with Fleet Operations and other engineering teams to resolve complex or recurring production issues. Perform root-cause analysis and post-incident reviews, ensuring corrective actions are documented, tracked, and completed. Maintain clear incident communications, runbooks, escalation procedures, and operational records. Support the team owned services in production environments and participate in the team's on-call rotation. Observability, Reliability, and Vendor Partnerships Own monitoring, dashboards, alerts, and operational KPIs for the team owned services using Prometheus and Grafana. Define reliability objectives and drive measurable reductions in incidents, escalations, and recurring support requests. Investigate hardware, firmware, and issues in partnership with internal engineering teams and external vendors. Manage vendor support cases involving BMCs, servers, power systems, and cooling infrastructure. Collect and provide diagnostic data, track issues through resolution, and validate vendor fixes before production rollout. Documentation and Continuous Improvement Create and maintain operational documentation, troubleshooting guides, escalation procedures, and service-support materials. Capture and share incident findings and operational knowledge across the team and partner teams. Identify repetitive operational tasks and develop more efficient, consistent remediation processes. Evaluate team processes using operational data and incident trends, and recommend improvements. Minimum Qualifications 5+ of experience in cloud operations, site reliability engineering (SRE), infrastructure operations, or a related technical field. Working knowledge of Kubernetes and at least one public cloud platform, such as AWS or GCP. Experience deploying and supporting containerized applications in Kubernetes environments. Experience with incident management practices, including incident response, escalation, and post-incident review processes. Experience using Prometheus, Grafana, and PromQL for monitoring, alerting, and troubleshooting. Strong knowledge of Linux system administration and internals and scripting. Experience troubleshooting complex issues across software services, operating systems, networks, and physical infrastructure. Experience participating in an on-call rotation supporting production services. Strong analytical and problem-solving skills, with a methodical approach to troubleshooting. Excellent written and verbal communication skills, particularly during high-impact incidents. Strong documentation skills and attention to detail. Preferred Qualifications Experience with server hardware, BMCs, Redfish, IPMI, or hardware-management services. Experience troubleshooting server provisioning, reboot, provisioning, or lifecycle-management failures. Familiarity with high-performance computing, GPU infrastructure, DPUs, or large-scale AI clusters. Experience working in data center environments, including server racks, power-distribution equipment, and cooling systems. Experience collaborating directly with hardware or firmware vendors to qualify and validate fixes. Bachelor's degree in computer science, engineering, or a related discipline-or equivalent practical experience. Understanding of Python or Golang. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You enjoy working close to the hardware and are curious about how GPUs, servers, and data centers fit together. You thrive in infrastructure environments where reliability, performance, and automation matter as much as features. You like collaborating across hardware, platform, and product teams to solve complex, ambiguous problems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $134,000 to $179,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace . click apply for full job details
09/23/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an experienced Senior Operations Engineer to join the MetalDev Operations team. Reporting to the Engineering Manager, this hands-on technical role focuses on service reliability, observability, and operational excellence. You will develop deep expertise in the Redfish-based services and automation tools used by frontline operations teams during data center bring-ups and in production environments. You will identify gaps in tooling and operational processes, partner with the engineering team to develop and validate fixes, and help ensure our automation operates reliably at scale. You will also submit change requests to hardware and firmware vendors, validate vendor-provided fixes, and continuously improve runbooks and on-call documentation to reduce recurring incidents and operational toil. In this role, you will work with cutting-edge infrastructure, including high-performance NVIDIA GPU servers, Cooling Distribution Units (CDUs), NVLink switches, and power shelves supporting GB200 and GB300 Vera Rubin NVL72 systems and custom in-house hardware. You will collaborate closely with the Hardware Engineering, Fleet Operations, and the Service Engineering teams. Approximately 80% of the role will focus on day-to-day operations, production support, and incident response. The remaining 20% will focus on improving operational processes, observability, documentation, and remediation capabilities and automation to prevent recurring incidents and increase service reliability. Key Responsibilities Triage and Troubleshooting Troubleshoot the team owned services, including initialization and reboot issues. Diagnose problems involving BMCs of, servers, DPUs, power shelves, Cooling Distribution Units (CDUs). Monitor fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Partner with Fleet Operations and other engineering teams to resolve complex or recurring production issues. Perform root-cause analysis and post-incident reviews, ensuring corrective actions are documented, tracked, and completed. Maintain clear incident communications, runbooks, escalation procedures, and operational records. Support the team owned services in production environments and participate in the team's on-call rotation. Observability, Reliability, and Vendor Partnerships Own monitoring, dashboards, alerts, and operational KPIs for the team owned services using Prometheus and Grafana. Define reliability objectives and drive measurable reductions in incidents, escalations, and recurring support requests. Investigate hardware, firmware, and issues in partnership with internal engineering teams and external vendors. Manage vendor support cases involving BMCs, servers, power systems, and cooling infrastructure. Collect and provide diagnostic data, track issues through resolution, and validate vendor fixes before production rollout. Documentation and Continuous Improvement Create and maintain operational documentation, troubleshooting guides, escalation procedures, and service-support materials. Capture and share incident findings and operational knowledge across the team and partner teams. Identify repetitive operational tasks and develop more efficient, consistent remediation processes. Evaluate team processes using operational data and incident trends, and recommend improvements. Minimum Qualifications 5+ of experience in cloud operations, site reliability engineering (SRE), infrastructure operations, or a related technical field. Working knowledge of Kubernetes and at least one public cloud platform, such as AWS or GCP. Experience deploying and supporting containerized applications in Kubernetes environments. Experience with incident management practices, including incident response, escalation, and post-incident review processes. Experience using Prometheus, Grafana, and PromQL for monitoring, alerting, and troubleshooting. Strong knowledge of Linux system administration and internals and scripting. Experience troubleshooting complex issues across software services, operating systems, networks, and physical infrastructure. Experience participating in an on-call rotation supporting production services. Strong analytical and problem-solving skills, with a methodical approach to troubleshooting. Excellent written and verbal communication skills, particularly during high-impact incidents. Strong documentation skills and attention to detail. Preferred Qualifications Experience with server hardware, BMCs, Redfish, IPMI, or hardware-management services. Experience troubleshooting server provisioning, reboot, provisioning, or lifecycle-management failures. Familiarity with high-performance computing, GPU infrastructure, DPUs, or large-scale AI clusters. Experience working in data center environments, including server racks, power-distribution equipment, and cooling systems. Experience collaborating directly with hardware or firmware vendors to qualify and validate fixes. Bachelor's degree in computer science, engineering, or a related discipline-or equivalent practical experience. Understanding of Python or Golang. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You enjoy working close to the hardware and are curious about how GPUs, servers, and data centers fit together. You thrive in infrastructure environments where reliability, performance, and automation matter as much as features. You like collaborating across hardware, platform, and product teams to solve complex, ambiguous problems. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $134,000 to $179,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace . click apply for full job details
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.
09/23/2026
Full time
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange ️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world's largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world's hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You'll Do (Role Expectations) Maintain high availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks in partnership with Engineering and Networking teams Lead full-cycle incident response by conducting cross-stack troubleshooting using low-level OS and network tools (strace, lsof, tcpdump, iostat, vmstat, gdb), maintain high availability across large-scale bare metal Linux /BSD fleets and Kubernetes clusters Automate infrastructure lifecycle management, service provisioning, configuration workflows, and release deployments using Ansible, Python, and Go; quantify operational toil and convert recurring manual work into durable, version-controlled, testable automation - tracking reduction as an engineering outcome Own end-to-end telemetry (metrics, logs, traces) using Prometheus and OpenTelemetry ecosystems; define and enforce SLOs/error budgets to reduce alert noise Perform architectural reviews, OS/kernel upgrades, capacity and performance tuning, strict CI/CD validation prior to production rollouts; embed operability standards (telemetry, rollback safety, SLO readiness) into service design from the start Who You Are (Success Profile) You thrive in ambiguity. You're comfortable building the path as you walk it, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution. You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback-knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We're Looking For (Minimum Qualifications) US Citizenship is required (due to the nature of assigned customers) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering operating high-scale, low-latency production platforms Proven ability to write and debug executable code live (Python, Go, or Bash) covering core logic/data structures, along with hands-on experience writing Ansible playbooks/tasks for infrastructure automation Deep knowledge of Linux OS internals and kernel troubleshooting (e.g., inodes, open file descriptors, process states, and analyzing df vs du storage discrepancies) Comprehensive understanding of networking protocols and packet-level analysis, including DNS resolution workflows, TLS handshakes, TCP/IP mechanics, and packet captures via tcpdump What Will Make You Stand Out (Preferred Qualifications) Hands-on experience operating and managing FreeBSD / BSD operating systems in production Proven expertise running, scaling, and troubleshooting Kubernetes clusters in high-traffic, low-latency environments and workflow orchestration platforms (Temporal or similar) Deep experience with Prometheus / OpenTelemetry ecosystems, or leveraging AI/ML frameworks/AIOps tools for automated root-cause analysis Zscaler's salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and relevant education or training. The base salary range listed for this full-time position excludes commission/ bonus/ equity (if applicable) + benefits. Base Pay Range $119,000-$170,000 USD At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.