Sager Electronics seeks a Field Application Engineer to support OEM and industrial customers with technical solutions in interconnect, power, and electromechanical components. You will collaborate with sales teams to identify design-in opportunities, recommend components, and provide on-site and virtual application support. Responsibilities include reviewing schematics, troubleshooting designs, conducting product demos, and liaising with manufacturers. In our customer-focused, collaborative culture, you'll grow your technical and supply chain expertise while driving design wins and contributing to innovative customer projects. Responsibilities Provide on-site and remote technical support to OEM and industrial customers Review schematics and designs to recommend appropriate components and solutions Collaborate with sales teams to identify and drive design-in opportunities Conduct product demonstrations, training, and technical presentations Troubleshoot application issues and coordinate with manufacturers for advanced support Support value-added and custom solutions for customer-specific requirements Document customer requirements and maintain opportunity data in CRM systems Stay current on product technologies and market trends to guide customers effectively Required Skills Electronics circuit design Power electronics Interconnect and electromechanical components Reading schematics and datasheets Technical sales support Design-in/application support Customer presentation skills CRM and sales tools Failure analysis and troubleshooting Supply chain and distribution familiarity
09/23/2026
Full time
Sager Electronics seeks a Field Application Engineer to support OEM and industrial customers with technical solutions in interconnect, power, and electromechanical components. You will collaborate with sales teams to identify design-in opportunities, recommend components, and provide on-site and virtual application support. Responsibilities include reviewing schematics, troubleshooting designs, conducting product demos, and liaising with manufacturers. In our customer-focused, collaborative culture, you'll grow your technical and supply chain expertise while driving design wins and contributing to innovative customer projects. Responsibilities Provide on-site and remote technical support to OEM and industrial customers Review schematics and designs to recommend appropriate components and solutions Collaborate with sales teams to identify and drive design-in opportunities Conduct product demonstrations, training, and technical presentations Troubleshoot application issues and coordinate with manufacturers for advanced support Support value-added and custom solutions for customer-specific requirements Document customer requirements and maintain opportunity data in CRM systems Stay current on product technologies and market trends to guide customers effectively Required Skills Electronics circuit design Power electronics Interconnect and electromechanical components Reading schematics and datasheets Technical sales support Design-in/application support Customer presentation skills CRM and sales tools Failure analysis and troubleshooting Supply chain and distribution familiarity
Sager Electronics is seeking a Field Application Engineer to support OEM and industrial customers across the wholesale electronic components market. In this customer-facing role, you'll provide design-in support, recommend interconnect, power, and electromechanical solutions, and resolve complex technical challenges. Partner with sales, suppliers, and engineering teams to deliver value-added services, on-site visits, and product demos. You'll leverage your electronics expertise, communication skills, and problem-solving abilities in a collaborative, growth-oriented, and service-driven culture with strong opportunities for development. Responsibilities Provide on-site and remote technical support to OEM and industrial customers. Recommend and design-in interconnect, power, and electromechanical component solutions. Collaborate with sales teams to develop and execute technical strategies for key accounts. Conduct product demos, technical presentations, and training for customers and internal teams. Analyze customer requirements, schematics, and specifications to propose optimal solutions. Support troubleshooting, root cause analysis, and resolution of application issues in the field. Interface with supplier engineering teams to stay current on products and technologies. Document customer interactions, opportunities, and designs within CRM/ERP systems. Contribute to continuous improvement of technical processes, tools, and support resources. Build long-term customer relationships aligned with Sager's service-driven culture. Required Skills Electronic circuit design and analysis Interconnect, power, and electromechanical components knowledge Reading schematics and technical documentation Customer-facing technical support Solution selling / design-in support Troubleshooting and failure analysis Technical presentations and demos CRM and ERP systems use Project and time management Microsoft Office and collaboration tools
09/23/2026
Full time
Sager Electronics is seeking a Field Application Engineer to support OEM and industrial customers across the wholesale electronic components market. In this customer-facing role, you'll provide design-in support, recommend interconnect, power, and electromechanical solutions, and resolve complex technical challenges. Partner with sales, suppliers, and engineering teams to deliver value-added services, on-site visits, and product demos. You'll leverage your electronics expertise, communication skills, and problem-solving abilities in a collaborative, growth-oriented, and service-driven culture with strong opportunities for development. Responsibilities Provide on-site and remote technical support to OEM and industrial customers. Recommend and design-in interconnect, power, and electromechanical component solutions. Collaborate with sales teams to develop and execute technical strategies for key accounts. Conduct product demos, technical presentations, and training for customers and internal teams. Analyze customer requirements, schematics, and specifications to propose optimal solutions. Support troubleshooting, root cause analysis, and resolution of application issues in the field. Interface with supplier engineering teams to stay current on products and technologies. Document customer interactions, opportunities, and designs within CRM/ERP systems. Contribute to continuous improvement of technical processes, tools, and support resources. Build long-term customer relationships aligned with Sager's service-driven culture. Required Skills Electronic circuit design and analysis Interconnect, power, and electromechanical components knowledge Reading schematics and technical documentation Customer-facing technical support Solution selling / design-in support Troubleshooting and failure analysis Technical presentations and demos CRM and ERP systems use Project and time management Microsoft Office and collaboration tools
Summary The Senior Data Integration Engineer is responsible for the design, development, and operation of the data integration and ingestion processes that deliver partner and internal data into our analytics environment. The role owns the flow of data from source acquisition through the curated data warehouse tables consumed by reporting platforms, operational systems, and business analysts. This is a senior, hands-on engineering position spanning the full integration lifecycle: acquiring data from a wide range of external partner and internal systems, validating and conditioning that data on arrival, transforming it into the structures that support analysis and operations, and operating those processes reliably in production. The role is concerned equally with building new integrations and with the continued performance, accuracy, and timeliness of those already in service. A central objective of the position is to advance reusable, well-instrumented integration patterns that shorten the time required to onboard new data sources and that improve the reliability and transparency of data delivery to the business. The Senior Data Integration Engineer will also contribute substantially to the planned modernization of our data platform, evaluating and recommending tooling, architecture, and migration approach for leadership consideration. The role sets technical direction and development standards for data integration work and collaborates closely with data architects, business analysts, stakeholders across the organization, and the technical contacts of our external data partners. Essential Duties and Responsibilities This list of duties and responsibilities is not all inclusive and may be expanded to include other duties and responsibilities as management may deem necessary from time to time. Design, develop, and maintain data integration processes that acquire data from partner and internal sources, including flat file transfers over SFTP, REST API endpoints, and direct database connections. Develop reusable, configuration-driven ingestion patterns that reduce the effort and elapsed time required to onboard new partner data feeds. Develop and maintain the T-SQL transformation logic that carries data from landing and staging layers through to the curated warehouse tables supporting reporting, operational systems, and analyst queries. Design ingestion processes to be idempotent and safely re-runnable, incorporating automated retry and restart behavior for failed executions. Implement automated validation and quarantine processes so that records failing business-defined quality rules are isolated, reported, and prevented from reaching downstream consumers. Implement data quality rules defined by the business, including schema validation, reconciliation, row count and threshold checks, and anomaly detection. Establish monitoring, logging, and alerting for pipeline execution state, data freshness, and load completion, and automate the communication of ingestion status to stakeholders. Diagnose and resolve production data incidents, determine root cause, coordinate remediation, and document preventive measures through runbooks and post-incident review. Contribute to dimensional data model design in collaboration with the data architect and senior team members. Establish and maintain version control, code review, and repeatable deployment practices for database and pipeline code. Define and uphold development standards for data integration work through code review and technical guidance. Evaluate and recommend tooling, architecture, and sequencing for the platform modernization effort for leadership consideration. Migrate established integration workflows to modernized patterns incrementally and without disruption to production operations. Maintain documentation of data feeds, dependencies, lineage, ownership, and escalation paths. Perform all work involving protected health information in accordance with HIPAA requirements, including least-privilege access, secure transmission and storage of partner data, and the exclusion of PHI from logs and non-production environments. Coordinate with partner technical contacts, as needed, to resolve file format, schema, and connectivity questions. Provide occasional off-hours support for critical data load failures or production support rotations as needed to support timely response to critical data issues. Support AI and machine learning initiatives by maintaining reliable, secure, and well-governed data pipelines and datasets used for model development, testing, deployment, monitoring, and ongoing performance evaluation. Maintain confidentiality of information processed & follow company policies and procedures. Qualifications Requires six (6) or more years of professional data engineering, data operations, data platform operations, or related experience. Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent professional experience preferred. Experience leading teams and developing supervisory staff preferred. To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed below are representative of the knowledge, skill, and/or ability required. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. Other qualifications include: Advanced T-SQL development skills, including: Set-based rewriting of row-by-row and cursor-based logic. MERGE, upsert, and slowly changing dimension load patterns. Window functions and complex analytic queries. Execution plan analysis, index strategy, statistics management, and resolution of performance issues such as parameter sniffing. Transaction management and structured error handling within stored procedures, including correct rollback behavior on partial failure. Demonstrated proficiency in Python for data engineering applications, including API-based data acquisition, file parsing and format handling, data validation, and the development of packaged, scheduled jobs. Familiarity with common data libraries such as requests and pandas is expected. Demonstrated experience acquiring and integrating data from heterogeneous sources, including delimited, fixed-width, JSON, and XML file formats; REST APIs requiring authentication, pagination, and rate-limit handling; and direct database connectivity. Proficiency with Git and collaborative development workflows, including branching, pull requests, and code review. A code-first development approach, with integration logic authored and maintained in T-SQL and Python under source control. Ability to analyze pipeline and query performance and to improve the reliability, scalability, and cost efficiency of data workloads. Strong written communication skills, with the ability to produce runbooks, technical documentation, and incident reports, and to convey the business impact of technical issues to non-technical stakeholders. Strong problem-solving skills, attention to detail, and demonstrated ownership of production systems. Experience with Microsoft Azure data services such as Azure Data Factory or Microsoft Fabric, or comparable cloud orchestration platforms. Experience migrating on-premises SQL Server integration workloads to a cloud platform. Experience with dimensional modeling and data warehouse design. Experience designing and rationalizing SQL Server Agent job dependencies and scheduling. Experience establishing version control, code review, and repeatable deployment practices for database code. Experience implementing CI/CD pipelines for database projects. Experience handling protected health information under HIPAA, or comparably regulated data under an equivalent framework preferred. Experience with healthcare or pharmacy data, including claims, eligibility, prescription, or delivery data preferred. Familiarity with data governance, metadata management, and data lineage practices. Physical Demands The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. (The phrases "occasionally," "regularly," and "frequently" correspond to the following definitions: "Occasionally" means up to 1/3 of working time, "regularly" means between 1/3 and 2/3 of working time, and "frequently" means 2/3 and more working time.) While performing the duties of this job, the employee is frequently required to sit; talk or hear; and use hands to handle, or touch objects or controls. The employee is regularly required to stand and walk. On occasion the incumbent may be required to stoop, bend or reach above the shoulders. The employee would rarely need to lift up to 25 pounds. Specific vision abilities required by this job include close vision, distance vision, color vision, peripheral vision, depth perception, and ability to adjust focus. Work Environment The position is a hybrid position with 2 - 3 days office presence in Milwaukee, WI required. The colleague may perform work-related travel on rare occasions (less than 10%) with an emphasis on travel for impact. . click apply for full job details
09/23/2026
Full time
Summary The Senior Data Integration Engineer is responsible for the design, development, and operation of the data integration and ingestion processes that deliver partner and internal data into our analytics environment. The role owns the flow of data from source acquisition through the curated data warehouse tables consumed by reporting platforms, operational systems, and business analysts. This is a senior, hands-on engineering position spanning the full integration lifecycle: acquiring data from a wide range of external partner and internal systems, validating and conditioning that data on arrival, transforming it into the structures that support analysis and operations, and operating those processes reliably in production. The role is concerned equally with building new integrations and with the continued performance, accuracy, and timeliness of those already in service. A central objective of the position is to advance reusable, well-instrumented integration patterns that shorten the time required to onboard new data sources and that improve the reliability and transparency of data delivery to the business. The Senior Data Integration Engineer will also contribute substantially to the planned modernization of our data platform, evaluating and recommending tooling, architecture, and migration approach for leadership consideration. The role sets technical direction and development standards for data integration work and collaborates closely with data architects, business analysts, stakeholders across the organization, and the technical contacts of our external data partners. Essential Duties and Responsibilities This list of duties and responsibilities is not all inclusive and may be expanded to include other duties and responsibilities as management may deem necessary from time to time. Design, develop, and maintain data integration processes that acquire data from partner and internal sources, including flat file transfers over SFTP, REST API endpoints, and direct database connections. Develop reusable, configuration-driven ingestion patterns that reduce the effort and elapsed time required to onboard new partner data feeds. Develop and maintain the T-SQL transformation logic that carries data from landing and staging layers through to the curated warehouse tables supporting reporting, operational systems, and analyst queries. Design ingestion processes to be idempotent and safely re-runnable, incorporating automated retry and restart behavior for failed executions. Implement automated validation and quarantine processes so that records failing business-defined quality rules are isolated, reported, and prevented from reaching downstream consumers. Implement data quality rules defined by the business, including schema validation, reconciliation, row count and threshold checks, and anomaly detection. Establish monitoring, logging, and alerting for pipeline execution state, data freshness, and load completion, and automate the communication of ingestion status to stakeholders. Diagnose and resolve production data incidents, determine root cause, coordinate remediation, and document preventive measures through runbooks and post-incident review. Contribute to dimensional data model design in collaboration with the data architect and senior team members. Establish and maintain version control, code review, and repeatable deployment practices for database and pipeline code. Define and uphold development standards for data integration work through code review and technical guidance. Evaluate and recommend tooling, architecture, and sequencing for the platform modernization effort for leadership consideration. Migrate established integration workflows to modernized patterns incrementally and without disruption to production operations. Maintain documentation of data feeds, dependencies, lineage, ownership, and escalation paths. Perform all work involving protected health information in accordance with HIPAA requirements, including least-privilege access, secure transmission and storage of partner data, and the exclusion of PHI from logs and non-production environments. Coordinate with partner technical contacts, as needed, to resolve file format, schema, and connectivity questions. Provide occasional off-hours support for critical data load failures or production support rotations as needed to support timely response to critical data issues. Support AI and machine learning initiatives by maintaining reliable, secure, and well-governed data pipelines and datasets used for model development, testing, deployment, monitoring, and ongoing performance evaluation. Maintain confidentiality of information processed & follow company policies and procedures. Qualifications Requires six (6) or more years of professional data engineering, data operations, data platform operations, or related experience. Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent professional experience preferred. Experience leading teams and developing supervisory staff preferred. To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed below are representative of the knowledge, skill, and/or ability required. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. Other qualifications include: Advanced T-SQL development skills, including: Set-based rewriting of row-by-row and cursor-based logic. MERGE, upsert, and slowly changing dimension load patterns. Window functions and complex analytic queries. Execution plan analysis, index strategy, statistics management, and resolution of performance issues such as parameter sniffing. Transaction management and structured error handling within stored procedures, including correct rollback behavior on partial failure. Demonstrated proficiency in Python for data engineering applications, including API-based data acquisition, file parsing and format handling, data validation, and the development of packaged, scheduled jobs. Familiarity with common data libraries such as requests and pandas is expected. Demonstrated experience acquiring and integrating data from heterogeneous sources, including delimited, fixed-width, JSON, and XML file formats; REST APIs requiring authentication, pagination, and rate-limit handling; and direct database connectivity. Proficiency with Git and collaborative development workflows, including branching, pull requests, and code review. A code-first development approach, with integration logic authored and maintained in T-SQL and Python under source control. Ability to analyze pipeline and query performance and to improve the reliability, scalability, and cost efficiency of data workloads. Strong written communication skills, with the ability to produce runbooks, technical documentation, and incident reports, and to convey the business impact of technical issues to non-technical stakeholders. Strong problem-solving skills, attention to detail, and demonstrated ownership of production systems. Experience with Microsoft Azure data services such as Azure Data Factory or Microsoft Fabric, or comparable cloud orchestration platforms. Experience migrating on-premises SQL Server integration workloads to a cloud platform. Experience with dimensional modeling and data warehouse design. Experience designing and rationalizing SQL Server Agent job dependencies and scheduling. Experience establishing version control, code review, and repeatable deployment practices for database code. Experience implementing CI/CD pipelines for database projects. Experience handling protected health information under HIPAA, or comparably regulated data under an equivalent framework preferred. Experience with healthcare or pharmacy data, including claims, eligibility, prescription, or delivery data preferred. Familiarity with data governance, metadata management, and data lineage practices. Physical Demands The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. (The phrases "occasionally," "regularly," and "frequently" correspond to the following definitions: "Occasionally" means up to 1/3 of working time, "regularly" means between 1/3 and 2/3 of working time, and "frequently" means 2/3 and more working time.) While performing the duties of this job, the employee is frequently required to sit; talk or hear; and use hands to handle, or touch objects or controls. The employee is regularly required to stand and walk. On occasion the incumbent may be required to stoop, bend or reach above the shoulders. The employee would rarely need to lift up to 25 pounds. Specific vision abilities required by this job include close vision, distance vision, color vision, peripheral vision, depth perception, and ability to adjust focus. Work Environment The position is a hybrid position with 2 - 3 days office presence in Milwaukee, WI required. The colleague may perform work-related travel on rare occasions (less than 10%) with an emphasis on travel for impact. . click apply for full job details
The work we do has an impact on millions of lives, and you can be a part of it. We help protect our customers against life's uncertainties. Regardless of where you work within the company, you'll be helping provide protection and peace of mind when our customers need it most. Protective is looking for a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, our Databricks lakehouse on Azure. You will lead the design of data products through the full medallion architecture - Bronze ingestion, Silver conformance, and Gold consumption - and be accountable for whether those products hold up for the consumers who depend on them. This is a hands-on technical leadership role, not a people-management or project-management role. You will still write and review production code daily. What you own is how the pod's data products are designed, modeled, contracted, and tested, and the standard the pod holds itself to. The Product Owner owns the backlog and the Scrum Master owns the sprint; you own the engineering. On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description. Key Responsibilities Design and modeling • Lead the design of data products end to end: what gets ingested, how it is cleansed and conformed, how it is modeled, and what the Gold layer looks like to the people querying it. • Own dimensional design - grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across the pod's products. • Set the pod's position on where logic belongs: what is cleaned in Silver, what is business logic in Gold, and what is a consumer's own concern. • Keep models as simple as the questions require, and push back on designs that will not hold. • Partner with ML engineering where the pod's Gold layer is the training or feature source for a model, so those datasets are contracted, versioned, and reproducible like any other consumer-facing product. Data contracts and consumer compatibility • Own the pod's ODCS data contracts as real interfaces: named owners, named consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy. • Make the compatibility call on every proposed contract change, and drive consumer notification when a change is genuinely breaking. • Represent the pod's contracts in cross-domain conversations, where one pod's Gold layer is another team's dependency. Standards, quality, and operations • Set and hold the pod's engineering standards for Python, SQL, dbt, testing, model structure, naming, and repository conventions - consistent with the platform's paved paths and Azure DevOps CI gates rather than in competition with them. • Lead code review. Be the reviewer who catches the modeling mistake, the missing test, and the change that will break a consumer, and who explains why so the pod learns it. • Ensure quality rules are enforced in tests and asset checks rather than asserted in documentation, and that pipeline health is observable without someone going to look - freshness, volume, latency, and cost instrumented, and alerting set against the SLAs and SLOs the pod's contracts commit to. • Own the pod's operational posture for its own pipelines: failure diagnosis, data-issue triage, backfills, cost and performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks someone other than the author can execute. • Keep the pod's delivery inside the control expectations of a regulated carrier: change management through pull request and pipeline, segregation of duties between authoring and deploying, least-privilege access, and audit evidence that falls out of CI/CD rather than being assembled for an auditor. • Develop reusable frameworks, templates, and patterns that raise the pod's consistency and delivery speed. Delivery leadership • Partner with the Product Owner and Scrum Master on decomposition and refinement: turn use cases and features into estimable, grounded stories with testable acceptance criteria and an identified target layer and repository. • Hold the Definition of Ready before the pod commits and the Definition of Done before the pod calls something finished - merged and approved code, passing CI and coverage gates, and evidence that the outcome is real. • Identify unknowns that need a spike rather than an estimate, and say so during planning rather than mid-sprint. Mentorship and collaboration • Grow the engineers on the pod through design review, pairing, and code review rather than by taking the hard work yourself. • Bring new engineers up to productive speed on the platform's conventions and tooling, and reduce single points of knowledge - no data product that only one person understands. • Work with the platform team on capability gaps: when the pod needs something the platform does not yet offer, raise it as a demand signal rather than building a private workaround. • Partner with the DataOps/MLOps Lead on the shared CI/CD, orchestration, and observability standards - adopt and strengthen the paved path rather than forking it, and be the pod's voice on what it is still missing. • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements. Qualifications Required Qualifications • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered. • 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through to published data products. • Strong hands-on Python and SQL, with the credibility to make design calls and the willingness to still write and review code. • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning. • Deep dimensional modeling experience - grain, keys, slowly changing dimensions, facts and dimensions, conformed dimensions. • Demonstrated technical leadership: setting standards, leading design, and raising other engineers' work, whether or not the role carried a lead title. • Experience owning data that other teams depend on, including handling breaking changes and production data incidents. • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD - Azure DevOps or comparable. • Experience setting observability and SLA/SLO expectations for data other teams depend on, and running the incident and communication path when they are missed. • Ability to explain trade-offs clearly to engineers, product owners, and business stakeholders, and to say no to a design that will not hold. Preferred Qualifications • Databricks certification (Data Engineer Professional or equivalent demonstrated depth). • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage. • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake. • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments. • Practical experience with data contracts, ODCS, or data-mesh style data product ownership. • Experience with a declarative Python ingestion framework such as dlt (dltHub) or comparable. • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar. • Familiarity with MLOps practice - MLflow, model registries, and model serving - sufficient to design data products that ML systems can depend on. • Azure and Azure DevOps. • Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations. • Experience introducing AI-assisted development into a team's normal workflow in a disciplined way. $109,500 - $167,833 a year Protective's targeted salary range for this position is $109,500 to $167,833. Actual salaries may vary depending on factors, including but not limited to, job location, skills, and experience. The range listed is just one component of Protective's total compensation package for employees. This position also offers additional incentive opportunities through an annual incentive based on individual and Company performance. Employee Benefits: We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits ( e.g. , paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health. Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective's platform to improve wellbeing while earning cash rewards. Eligibility for certain benefits may vary by position in accordance with the terms of the Company's benefit plans. Accommodations for Applicants with a Disability: If you require an accommodation to complete the application and recruitment process due to a disability, please email eric . click apply for full job details
09/23/2026
Full time
The work we do has an impact on millions of lives, and you can be a part of it. We help protect our customers against life's uncertainties. Regardless of where you work within the company, you'll be helping provide protection and peace of mind when our customers need it most. Protective is looking for a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, our Databricks lakehouse on Azure. You will lead the design of data products through the full medallion architecture - Bronze ingestion, Silver conformance, and Gold consumption - and be accountable for whether those products hold up for the consumers who depend on them. This is a hands-on technical leadership role, not a people-management or project-management role. You will still write and review production code daily. What you own is how the pod's data products are designed, modeled, contracted, and tested, and the standard the pod holds itself to. The Product Owner owns the backlog and the Scrum Master owns the sprint; you own the engineering. On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description. Key Responsibilities Design and modeling • Lead the design of data products end to end: what gets ingested, how it is cleansed and conformed, how it is modeled, and what the Gold layer looks like to the people querying it. • Own dimensional design - grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across the pod's products. • Set the pod's position on where logic belongs: what is cleaned in Silver, what is business logic in Gold, and what is a consumer's own concern. • Keep models as simple as the questions require, and push back on designs that will not hold. • Partner with ML engineering where the pod's Gold layer is the training or feature source for a model, so those datasets are contracted, versioned, and reproducible like any other consumer-facing product. Data contracts and consumer compatibility • Own the pod's ODCS data contracts as real interfaces: named owners, named consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy. • Make the compatibility call on every proposed contract change, and drive consumer notification when a change is genuinely breaking. • Represent the pod's contracts in cross-domain conversations, where one pod's Gold layer is another team's dependency. Standards, quality, and operations • Set and hold the pod's engineering standards for Python, SQL, dbt, testing, model structure, naming, and repository conventions - consistent with the platform's paved paths and Azure DevOps CI gates rather than in competition with them. • Lead code review. Be the reviewer who catches the modeling mistake, the missing test, and the change that will break a consumer, and who explains why so the pod learns it. • Ensure quality rules are enforced in tests and asset checks rather than asserted in documentation, and that pipeline health is observable without someone going to look - freshness, volume, latency, and cost instrumented, and alerting set against the SLAs and SLOs the pod's contracts commit to. • Own the pod's operational posture for its own pipelines: failure diagnosis, data-issue triage, backfills, cost and performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks someone other than the author can execute. • Keep the pod's delivery inside the control expectations of a regulated carrier: change management through pull request and pipeline, segregation of duties between authoring and deploying, least-privilege access, and audit evidence that falls out of CI/CD rather than being assembled for an auditor. • Develop reusable frameworks, templates, and patterns that raise the pod's consistency and delivery speed. Delivery leadership • Partner with the Product Owner and Scrum Master on decomposition and refinement: turn use cases and features into estimable, grounded stories with testable acceptance criteria and an identified target layer and repository. • Hold the Definition of Ready before the pod commits and the Definition of Done before the pod calls something finished - merged and approved code, passing CI and coverage gates, and evidence that the outcome is real. • Identify unknowns that need a spike rather than an estimate, and say so during planning rather than mid-sprint. Mentorship and collaboration • Grow the engineers on the pod through design review, pairing, and code review rather than by taking the hard work yourself. • Bring new engineers up to productive speed on the platform's conventions and tooling, and reduce single points of knowledge - no data product that only one person understands. • Work with the platform team on capability gaps: when the pod needs something the platform does not yet offer, raise it as a demand signal rather than building a private workaround. • Partner with the DataOps/MLOps Lead on the shared CI/CD, orchestration, and observability standards - adopt and strengthen the paved path rather than forking it, and be the pod's voice on what it is still missing. • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements. Qualifications Required Qualifications • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered. • 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through to published data products. • Strong hands-on Python and SQL, with the credibility to make design calls and the willingness to still write and review code. • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning. • Deep dimensional modeling experience - grain, keys, slowly changing dimensions, facts and dimensions, conformed dimensions. • Demonstrated technical leadership: setting standards, leading design, and raising other engineers' work, whether or not the role carried a lead title. • Experience owning data that other teams depend on, including handling breaking changes and production data incidents. • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD - Azure DevOps or comparable. • Experience setting observability and SLA/SLO expectations for data other teams depend on, and running the incident and communication path when they are missed. • Ability to explain trade-offs clearly to engineers, product owners, and business stakeholders, and to say no to a design that will not hold. Preferred Qualifications • Databricks certification (Data Engineer Professional or equivalent demonstrated depth). • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage. • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake. • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments. • Practical experience with data contracts, ODCS, or data-mesh style data product ownership. • Experience with a declarative Python ingestion framework such as dlt (dltHub) or comparable. • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar. • Familiarity with MLOps practice - MLflow, model registries, and model serving - sufficient to design data products that ML systems can depend on. • Azure and Azure DevOps. • Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations. • Experience introducing AI-assisted development into a team's normal workflow in a disciplined way. $109,500 - $167,833 a year Protective's targeted salary range for this position is $109,500 to $167,833. Actual salaries may vary depending on factors, including but not limited to, job location, skills, and experience. The range listed is just one component of Protective's total compensation package for employees. This position also offers additional incentive opportunities through an annual incentive based on individual and Company performance. Employee Benefits: We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits ( e.g. , paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health. Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective's platform to improve wellbeing while earning cash rewards. Eligibility for certain benefits may vary by position in accordance with the terms of the Company's benefit plans. Accommodations for Applicants with a Disability: If you require an accommodation to complete the application and recruitment process due to a disability, please email eric . click apply for full job details
The work we do has an impact on millions of lives, and you can be a part of it. We help protect our customers against life's uncertainties. Regardless of where you work within the company, you'll be helping provide protection and peace of mind when our customers need it most. Protective is looking for a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, our Databricks lakehouse on Azure. You will lead the design of data products through the full medallion architecture - Bronze ingestion, Silver conformance, and Gold consumption - and be accountable for whether those products hold up for the consumers who depend on them. This is a hands-on technical leadership role, not a people-management or project-management role. You will still write and review production code daily. What you own is how the pod's data products are designed, modeled, contracted, and tested, and the standard the pod holds itself to. The Product Owner owns the backlog and the Scrum Master owns the sprint; you own the engineering. On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description. Key Responsibilities Design and modeling • Lead the design of data products end to end: what gets ingested, how it is cleansed and conformed, how it is modeled, and what the Gold layer looks like to the people querying it. • Own dimensional design - grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across the pod's products. • Set the pod's position on where logic belongs: what is cleaned in Silver, what is business logic in Gold, and what is a consumer's own concern. • Keep models as simple as the questions require, and push back on designs that will not hold. • Partner with ML engineering where the pod's Gold layer is the training or feature source for a model, so those datasets are contracted, versioned, and reproducible like any other consumer-facing product. Data contracts and consumer compatibility • Own the pod's ODCS data contracts as real interfaces: named owners, named consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy. • Make the compatibility call on every proposed contract change, and drive consumer notification when a change is genuinely breaking. • Represent the pod's contracts in cross-domain conversations, where one pod's Gold layer is another team's dependency. Standards, quality, and operations • Set and hold the pod's engineering standards for Python, SQL, dbt, testing, model structure, naming, and repository conventions - consistent with the platform's paved paths and Azure DevOps CI gates rather than in competition with them. • Lead code review. Be the reviewer who catches the modeling mistake, the missing test, and the change that will break a consumer, and who explains why so the pod learns it. • Ensure quality rules are enforced in tests and asset checks rather than asserted in documentation, and that pipeline health is observable without someone going to look - freshness, volume, latency, and cost instrumented, and alerting set against the SLAs and SLOs the pod's contracts commit to. • Own the pod's operational posture for its own pipelines: failure diagnosis, data-issue triage, backfills, cost and performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks someone other than the author can execute. • Keep the pod's delivery inside the control expectations of a regulated carrier: change management through pull request and pipeline, segregation of duties between authoring and deploying, least-privilege access, and audit evidence that falls out of CI/CD rather than being assembled for an auditor. • Develop reusable frameworks, templates, and patterns that raise the pod's consistency and delivery speed. Delivery leadership • Partner with the Product Owner and Scrum Master on decomposition and refinement: turn use cases and features into estimable, grounded stories with testable acceptance criteria and an identified target layer and repository. • Hold the Definition of Ready before the pod commits and the Definition of Done before the pod calls something finished - merged and approved code, passing CI and coverage gates, and evidence that the outcome is real. • Identify unknowns that need a spike rather than an estimate, and say so during planning rather than mid-sprint. Mentorship and collaboration • Grow the engineers on the pod through design review, pairing, and code review rather than by taking the hard work yourself. • Bring new engineers up to productive speed on the platform's conventions and tooling, and reduce single points of knowledge - no data product that only one person understands. • Work with the platform team on capability gaps: when the pod needs something the platform does not yet offer, raise it as a demand signal rather than building a private workaround. • Partner with the DataOps/MLOps Lead on the shared CI/CD, orchestration, and observability standards - adopt and strengthen the paved path rather than forking it, and be the pod's voice on what it is still missing. • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements. Qualifications Required Qualifications • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered. • 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through to published data products. • Strong hands-on Python and SQL, with the credibility to make design calls and the willingness to still write and review code. • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning. • Deep dimensional modeling experience - grain, keys, slowly changing dimensions, facts and dimensions, conformed dimensions. • Demonstrated technical leadership: setting standards, leading design, and raising other engineers' work, whether or not the role carried a lead title. • Experience owning data that other teams depend on, including handling breaking changes and production data incidents. • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD - Azure DevOps or comparable. • Experience setting observability and SLA/SLO expectations for data other teams depend on, and running the incident and communication path when they are missed. • Ability to explain trade-offs clearly to engineers, product owners, and business stakeholders, and to say no to a design that will not hold. Preferred Qualifications • Databricks certification (Data Engineer Professional or equivalent demonstrated depth). • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage. • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake. • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments. • Practical experience with data contracts, ODCS, or data-mesh style data product ownership. • Experience with a declarative Python ingestion framework such as dlt (dltHub) or comparable. • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar. • Familiarity with MLOps practice - MLflow, model registries, and model serving - sufficient to design data products that ML systems can depend on. • Azure and Azure DevOps. • Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations. • Experience introducing AI-assisted development into a team's normal workflow in a disciplined way. $109,500 - $167,833 a year Protective's targeted salary range for this position is $109,500 to $167,833. Actual salaries may vary depending on factors, including but not limited to, job location, skills, and experience. The range listed is just one component of Protective's total compensation package for employees. This position also offers additional incentive opportunities through an annual incentive based on individual and Company performance. Employee Benefits: We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits ( e.g. , paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health. Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective's platform to improve wellbeing while earning cash rewards. Eligibility for certain benefits may vary by position in accordance with the terms of the Company's benefit plans. Accommodations for Applicants with a Disability: If you require an accommodation to complete the application and recruitment process due to a disability, please email eric . click apply for full job details
09/23/2026
Full time
The work we do has an impact on millions of lives, and you can be a part of it. We help protect our customers against life's uncertainties. Regardless of where you work within the company, you'll be helping provide protection and peace of mind when our customers need it most. Protective is looking for a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, our Databricks lakehouse on Azure. You will lead the design of data products through the full medallion architecture - Bronze ingestion, Silver conformance, and Gold consumption - and be accountable for whether those products hold up for the consumers who depend on them. This is a hands-on technical leadership role, not a people-management or project-management role. You will still write and review production code daily. What you own is how the pod's data products are designed, modeled, contracted, and tested, and the standard the pod holds itself to. The Product Owner owns the backlog and the Scrum Master owns the sprint; you own the engineering. On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description. Key Responsibilities Design and modeling • Lead the design of data products end to end: what gets ingested, how it is cleansed and conformed, how it is modeled, and what the Gold layer looks like to the people querying it. • Own dimensional design - grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across the pod's products. • Set the pod's position on where logic belongs: what is cleaned in Silver, what is business logic in Gold, and what is a consumer's own concern. • Keep models as simple as the questions require, and push back on designs that will not hold. • Partner with ML engineering where the pod's Gold layer is the training or feature source for a model, so those datasets are contracted, versioned, and reproducible like any other consumer-facing product. Data contracts and consumer compatibility • Own the pod's ODCS data contracts as real interfaces: named owners, named consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy. • Make the compatibility call on every proposed contract change, and drive consumer notification when a change is genuinely breaking. • Represent the pod's contracts in cross-domain conversations, where one pod's Gold layer is another team's dependency. Standards, quality, and operations • Set and hold the pod's engineering standards for Python, SQL, dbt, testing, model structure, naming, and repository conventions - consistent with the platform's paved paths and Azure DevOps CI gates rather than in competition with them. • Lead code review. Be the reviewer who catches the modeling mistake, the missing test, and the change that will break a consumer, and who explains why so the pod learns it. • Ensure quality rules are enforced in tests and asset checks rather than asserted in documentation, and that pipeline health is observable without someone going to look - freshness, volume, latency, and cost instrumented, and alerting set against the SLAs and SLOs the pod's contracts commit to. • Own the pod's operational posture for its own pipelines: failure diagnosis, data-issue triage, backfills, cost and performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks someone other than the author can execute. • Keep the pod's delivery inside the control expectations of a regulated carrier: change management through pull request and pipeline, segregation of duties between authoring and deploying, least-privilege access, and audit evidence that falls out of CI/CD rather than being assembled for an auditor. • Develop reusable frameworks, templates, and patterns that raise the pod's consistency and delivery speed. Delivery leadership • Partner with the Product Owner and Scrum Master on decomposition and refinement: turn use cases and features into estimable, grounded stories with testable acceptance criteria and an identified target layer and repository. • Hold the Definition of Ready before the pod commits and the Definition of Done before the pod calls something finished - merged and approved code, passing CI and coverage gates, and evidence that the outcome is real. • Identify unknowns that need a spike rather than an estimate, and say so during planning rather than mid-sprint. Mentorship and collaboration • Grow the engineers on the pod through design review, pairing, and code review rather than by taking the hard work yourself. • Bring new engineers up to productive speed on the platform's conventions and tooling, and reduce single points of knowledge - no data product that only one person understands. • Work with the platform team on capability gaps: when the pod needs something the platform does not yet offer, raise it as a demand signal rather than building a private workaround. • Partner with the DataOps/MLOps Lead on the shared CI/CD, orchestration, and observability standards - adopt and strengthen the paved path rather than forking it, and be the pod's voice on what it is still missing. • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements. Qualifications Required Qualifications • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered. • 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through to published data products. • Strong hands-on Python and SQL, with the credibility to make design calls and the willingness to still write and review code. • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning. • Deep dimensional modeling experience - grain, keys, slowly changing dimensions, facts and dimensions, conformed dimensions. • Demonstrated technical leadership: setting standards, leading design, and raising other engineers' work, whether or not the role carried a lead title. • Experience owning data that other teams depend on, including handling breaking changes and production data incidents. • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD - Azure DevOps or comparable. • Experience setting observability and SLA/SLO expectations for data other teams depend on, and running the incident and communication path when they are missed. • Ability to explain trade-offs clearly to engineers, product owners, and business stakeholders, and to say no to a design that will not hold. Preferred Qualifications • Databricks certification (Data Engineer Professional or equivalent demonstrated depth). • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage. • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake. • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments. • Practical experience with data contracts, ODCS, or data-mesh style data product ownership. • Experience with a declarative Python ingestion framework such as dlt (dltHub) or comparable. • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar. • Familiarity with MLOps practice - MLflow, model registries, and model serving - sufficient to design data products that ML systems can depend on. • Azure and Azure DevOps. • Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations. • Experience introducing AI-assisted development into a team's normal workflow in a disciplined way. $109,500 - $167,833 a year Protective's targeted salary range for this position is $109,500 to $167,833. Actual salaries may vary depending on factors, including but not limited to, job location, skills, and experience. The range listed is just one component of Protective's total compensation package for employees. This position also offers additional incentive opportunities through an annual incentive based on individual and Company performance. Employee Benefits: We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits ( e.g. , paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health. Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective's platform to improve wellbeing while earning cash rewards. Eligibility for certain benefits may vary by position in accordance with the terms of the Company's benefit plans. Accommodations for Applicants with a Disability: If you require an accommodation to complete the application and recruitment process due to a disability, please email eric . click apply for full job details
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to lead pre-silicon verification of the Beowulf SoC, with focus on the Compute Subsystem (CSS), DDR memory subsystem, and Fabric NoC. This role will drive coverage, coherency, memory traffic, connectivity, error handling, and bring-up features critical to silicon success. This role is hybrid, based out of Boston, MA; Toronto, ON; or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Deeply curious about SoC architecture, compute systems, DDR behavior, and Fabric NoC/interconnect verification. Expert in UVM, SystemVerilog, coverage-driven verification, assertions, and subsystem-level debug. Experienced in verifying compute subsystems, DDR controllers and PHY-facing logic, NoC/interconnect protocols, coherency, ordering, and data movement. Comfortable with reset, power management, error handling, performance, and high-concurrency system scenarios. Proactive, detail-oriented, and effective in cross-functional technical discussions. Familiar with Python, C/C++, Tcl, CocoTB, or similar verification automation tools. What We Need Develop and own scalable verification environments for Beowulf SoC CSS, DDR, and Fabric NoC across subsystem and full-SoC levels. Write, refine, and execute test scenarios for compute operation, memory initialization and traffic, coherency, routing, ordering, QoS, power management, error handling, and data movement. Analyze coverage gaps, debug failures, drive root-cause analysis, and collaborate closely with architecture, RTL, firmware, emulation, and system teams. Drive verification planning, regression quality, coverage closure, and signoff for major SoC features. Automate verification flows using scripting, reusable infrastructure, and AI productivity tools. Mentor engineers and contribute to verification methodology and design-for-verification improvements. What You Will Learn In-depth SoC verification across compute, DDR, coherency, and Fabric NoC using modern workflows, tooling, and scripting. Integration of pre-silicon SoC verification with emulation, silicon bring-up, and platform validation strategies. How AI-driven automation reshapes modern DV workflows. Exposure to high-performance, system-level verification in advanced AI SoC designs. Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made. Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer. This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.
09/23/2026
Full time
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to lead pre-silicon verification of the Beowulf SoC, with focus on the Compute Subsystem (CSS), DDR memory subsystem, and Fabric NoC. This role will drive coverage, coherency, memory traffic, connectivity, error handling, and bring-up features critical to silicon success. This role is hybrid, based out of Boston, MA; Toronto, ON; or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Deeply curious about SoC architecture, compute systems, DDR behavior, and Fabric NoC/interconnect verification. Expert in UVM, SystemVerilog, coverage-driven verification, assertions, and subsystem-level debug. Experienced in verifying compute subsystems, DDR controllers and PHY-facing logic, NoC/interconnect protocols, coherency, ordering, and data movement. Comfortable with reset, power management, error handling, performance, and high-concurrency system scenarios. Proactive, detail-oriented, and effective in cross-functional technical discussions. Familiar with Python, C/C++, Tcl, CocoTB, or similar verification automation tools. What We Need Develop and own scalable verification environments for Beowulf SoC CSS, DDR, and Fabric NoC across subsystem and full-SoC levels. Write, refine, and execute test scenarios for compute operation, memory initialization and traffic, coherency, routing, ordering, QoS, power management, error handling, and data movement. Analyze coverage gaps, debug failures, drive root-cause analysis, and collaborate closely with architecture, RTL, firmware, emulation, and system teams. Drive verification planning, regression quality, coverage closure, and signoff for major SoC features. Automate verification flows using scripting, reusable infrastructure, and AI productivity tools. Mentor engineers and contribute to verification methodology and design-for-verification improvements. What You Will Learn In-depth SoC verification across compute, DDR, coherency, and Fabric NoC using modern workflows, tooling, and scripting. Integration of pre-silicon SoC verification with emulation, silicon bring-up, and platform validation strategies. How AI-driven automation reshapes modern DV workflows. Exposure to high-performance, system-level verification in advanced AI SoC designs. Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made. Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer. This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.
Job Title: Sr. Associate, Integration/Test Engineer Job Code: 44513 Job Location: Bristol, PA Job Schedule: 9/80- employees work 9 out of 14 days- totaling 80 hours- and have every other Friday off Job Description: L3Harris is looking for a Sr. Associate, Integration/Test Engineer with working knowledge of Telemetry and Radio Frequency (T&RF) technologies. Responsible for performing the integration and test of T&RF systems, ensuring electrical and physical compatibility to meet program technical, schedule and cost objectives. Integration responsibilities include development and implementation of integration plans and procedures for hardware and software. Testing responsibilities may include analyzing requirements for testability issues. Develop and implement both hardware and software system level test programs, plans, specifications, and procedures. Plan and lead test working groups, test readiness reviews. Conduct testability and producibility analyses and reviews. Essential Functions: Develop test plans and procedures based on design and program requirements. Develop, verify and validate test equipment solutions to support engineering development and production test. Responsible for supporting test technicians and engineering team with respect to troubleshooting test failures of the device under test (DUT), executing tests, proper use of electronic test equipment, or the associated electronic test equipment. Responsible for supporting Failure Review Board (FRB) and Root Cause Corrective Action (RCCA) analysis actions. Responsible for proactively supporting customer deliverable production goals. Communicates within and outside of own function to report status, roadblocks, improvement areas, support needs. Qualifications: Bachelor's Degree and minimum 2 years of prior relevant experience. Graduate Degree. In lieu of a degree, minimum of 6 years of prior related experience. Ability to obtain secret clearance. Preferred Additional Skills: Experience troubleshooting electronic hardware to the electrical component level. Experience in RF or Digital engineering troubleshooting techniques, tools, and process. Experience with high-speed digital and RF test equipment (network analyzers, spectrum analyzers, power meters, oscilloscope, logic analyzers, etc.). Highly motivated, self-starter and willing to learn. Excellent organizational skills. Strong analytical and problem-solving skills. Excellent verbal and written communication skills in a technical information environment. Prior experience in Telemetry and RF or related industry. Previous experience supporting DoD/government customers. Familiarity with LabVIEW/MATLAB/other tools for instrument control. Solid understanding of RF and electronic testing tools, equipment and techniques using high-speed digital and RF test equipment (network analyzers, spectrum analyzers, power meters, oscilloscope, logic analyzers, etc.). Supporting manufacturing production of electronic devices. Experience with OSHA requirements and ESD best practices. In compliance with pay transparency requirements, the salary range for this role in California, Massachusetts, New Jersey, Washington, and the Greater D.C, Denver, or NYC areas is $88,500 - $164,500. This is not a guarantee of compensation or salary, as final offer amount may vary based on factors including but not limited to experience and geographic location. L3Harris also offers a variety of benefits, including health and disability insurance, 401(k) match, flexible spending accounts, EAP, education assistance, parental leave, paid time off, and company-paid holidays. The specific programs and options available to an employee may vary depending on date of hire, schedule type, and the applicability of collective bargaining agreements
09/23/2026
Full time
Job Title: Sr. Associate, Integration/Test Engineer Job Code: 44513 Job Location: Bristol, PA Job Schedule: 9/80- employees work 9 out of 14 days- totaling 80 hours- and have every other Friday off Job Description: L3Harris is looking for a Sr. Associate, Integration/Test Engineer with working knowledge of Telemetry and Radio Frequency (T&RF) technologies. Responsible for performing the integration and test of T&RF systems, ensuring electrical and physical compatibility to meet program technical, schedule and cost objectives. Integration responsibilities include development and implementation of integration plans and procedures for hardware and software. Testing responsibilities may include analyzing requirements for testability issues. Develop and implement both hardware and software system level test programs, plans, specifications, and procedures. Plan and lead test working groups, test readiness reviews. Conduct testability and producibility analyses and reviews. Essential Functions: Develop test plans and procedures based on design and program requirements. Develop, verify and validate test equipment solutions to support engineering development and production test. Responsible for supporting test technicians and engineering team with respect to troubleshooting test failures of the device under test (DUT), executing tests, proper use of electronic test equipment, or the associated electronic test equipment. Responsible for supporting Failure Review Board (FRB) and Root Cause Corrective Action (RCCA) analysis actions. Responsible for proactively supporting customer deliverable production goals. Communicates within and outside of own function to report status, roadblocks, improvement areas, support needs. Qualifications: Bachelor's Degree and minimum 2 years of prior relevant experience. Graduate Degree. In lieu of a degree, minimum of 6 years of prior related experience. Ability to obtain secret clearance. Preferred Additional Skills: Experience troubleshooting electronic hardware to the electrical component level. Experience in RF or Digital engineering troubleshooting techniques, tools, and process. Experience with high-speed digital and RF test equipment (network analyzers, spectrum analyzers, power meters, oscilloscope, logic analyzers, etc.). Highly motivated, self-starter and willing to learn. Excellent organizational skills. Strong analytical and problem-solving skills. Excellent verbal and written communication skills in a technical information environment. Prior experience in Telemetry and RF or related industry. Previous experience supporting DoD/government customers. Familiarity with LabVIEW/MATLAB/other tools for instrument control. Solid understanding of RF and electronic testing tools, equipment and techniques using high-speed digital and RF test equipment (network analyzers, spectrum analyzers, power meters, oscilloscope, logic analyzers, etc.). Supporting manufacturing production of electronic devices. Experience with OSHA requirements and ESD best practices. In compliance with pay transparency requirements, the salary range for this role in California, Massachusetts, New Jersey, Washington, and the Greater D.C, Denver, or NYC areas is $88,500 - $164,500. This is not a guarantee of compensation or salary, as final offer amount may vary based on factors including but not limited to experience and geographic location. L3Harris also offers a variety of benefits, including health and disability insurance, 401(k) match, flexible spending accounts, EAP, education assistance, parental leave, paid time off, and company-paid holidays. The specific programs and options available to an employee may vary depending on date of hire, schedule type, and the applicability of collective bargaining agreements
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to lead pre-silicon verification of the Beowulf SoC, with focus on the Compute Subsystem (CSS), DDR memory subsystem, and Fabric NoC. This role will drive coverage, coherency, memory traffic, connectivity, error handling, and bring-up features critical to silicon success. This role is hybrid, based out of Boston, MA; Toronto, ON; or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Deeply curious about SoC architecture, compute systems, DDR behavior, and Fabric NoC/interconnect verification. Expert in UVM, SystemVerilog, coverage-driven verification, assertions, and subsystem-level debug. Experienced in verifying compute subsystems, DDR controllers and PHY-facing logic, NoC/interconnect protocols, coherency, ordering, and data movement. Comfortable with reset, power management, error handling, performance, and high-concurrency system scenarios. Proactive, detail-oriented, and effective in cross-functional technical discussions. Familiar with Python, C/C++, Tcl, CocoTB, or similar verification automation tools. What We Need Develop and own scalable verification environments for Beowulf SoC CSS, DDR, and Fabric NoC across subsystem and full-SoC levels. Write, refine, and execute test scenarios for compute operation, memory initialization and traffic, coherency, routing, ordering, QoS, power management, error handling, and data movement. Analyze coverage gaps, debug failures, drive root-cause analysis, and collaborate closely with architecture, RTL, firmware, emulation, and system teams. Drive verification planning, regression quality, coverage closure, and signoff for major SoC features. Automate verification flows using scripting, reusable infrastructure, and AI productivity tools. Mentor engineers and contribute to verification methodology and design-for-verification improvements. What You Will Learn In-depth SoC verification across compute, DDR, coherency, and Fabric NoC using modern workflows, tooling, and scripting. Integration of pre-silicon SoC verification with emulation, silicon bring-up, and platform validation strategies. How AI-driven automation reshapes modern DV workflows. Exposure to high-performance, system-level verification in advanced AI SoC designs. Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made. Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer. This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.
09/23/2026
Full time
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to lead pre-silicon verification of the Beowulf SoC, with focus on the Compute Subsystem (CSS), DDR memory subsystem, and Fabric NoC. This role will drive coverage, coherency, memory traffic, connectivity, error handling, and bring-up features critical to silicon success. This role is hybrid, based out of Boston, MA; Toronto, ON; or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Deeply curious about SoC architecture, compute systems, DDR behavior, and Fabric NoC/interconnect verification. Expert in UVM, SystemVerilog, coverage-driven verification, assertions, and subsystem-level debug. Experienced in verifying compute subsystems, DDR controllers and PHY-facing logic, NoC/interconnect protocols, coherency, ordering, and data movement. Comfortable with reset, power management, error handling, performance, and high-concurrency system scenarios. Proactive, detail-oriented, and effective in cross-functional technical discussions. Familiar with Python, C/C++, Tcl, CocoTB, or similar verification automation tools. What We Need Develop and own scalable verification environments for Beowulf SoC CSS, DDR, and Fabric NoC across subsystem and full-SoC levels. Write, refine, and execute test scenarios for compute operation, memory initialization and traffic, coherency, routing, ordering, QoS, power management, error handling, and data movement. Analyze coverage gaps, debug failures, drive root-cause analysis, and collaborate closely with architecture, RTL, firmware, emulation, and system teams. Drive verification planning, regression quality, coverage closure, and signoff for major SoC features. Automate verification flows using scripting, reusable infrastructure, and AI productivity tools. Mentor engineers and contribute to verification methodology and design-for-verification improvements. What You Will Learn In-depth SoC verification across compute, DDR, coherency, and Fabric NoC using modern workflows, tooling, and scripting. Integration of pre-silicon SoC verification with emulation, silicon bring-up, and platform validation strategies. How AI-driven automation reshapes modern DV workflows. Exposure to high-performance, system-level verification in advanced AI SoC designs. Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made. Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer. This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.
The work we do has an impact on millions of lives, and you can be a part of it. We help protect our customers against life's uncertainties. Regardless of where you work within the company, you'll be helping provide protection and peace of mind when our customers need it most. Protective is looking for a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, our Databricks lakehouse on Azure. You will lead the design of data products through the full medallion architecture - Bronze ingestion, Silver conformance, and Gold consumption - and be accountable for whether those products hold up for the consumers who depend on them. This is a hands-on technical leadership role, not a people-management or project-management role. You will still write and review production code daily. What you own is how the pod's data products are designed, modeled, contracted, and tested, and the standard the pod holds itself to. The Product Owner owns the backlog and the Scrum Master owns the sprint; you own the engineering. On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description. Key Responsibilities Design and modeling • Lead the design of data products end to end: what gets ingested, how it is cleansed and conformed, how it is modeled, and what the Gold layer looks like to the people querying it. • Own dimensional design - grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across the pod's products. • Set the pod's position on where logic belongs: what is cleaned in Silver, what is business logic in Gold, and what is a consumer's own concern. • Keep models as simple as the questions require, and push back on designs that will not hold. • Partner with ML engineering where the pod's Gold layer is the training or feature source for a model, so those datasets are contracted, versioned, and reproducible like any other consumer-facing product. Data contracts and consumer compatibility • Own the pod's ODCS data contracts as real interfaces: named owners, named consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy. • Make the compatibility call on every proposed contract change, and drive consumer notification when a change is genuinely breaking. • Represent the pod's contracts in cross-domain conversations, where one pod's Gold layer is another team's dependency. Standards, quality, and operations • Set and hold the pod's engineering standards for Python, SQL, dbt, testing, model structure, naming, and repository conventions - consistent with the platform's paved paths and Azure DevOps CI gates rather than in competition with them. • Lead code review. Be the reviewer who catches the modeling mistake, the missing test, and the change that will break a consumer, and who explains why so the pod learns it. • Ensure quality rules are enforced in tests and asset checks rather than asserted in documentation, and that pipeline health is observable without someone going to look - freshness, volume, latency, and cost instrumented, and alerting set against the SLAs and SLOs the pod's contracts commit to. • Own the pod's operational posture for its own pipelines: failure diagnosis, data-issue triage, backfills, cost and performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks someone other than the author can execute. • Keep the pod's delivery inside the control expectations of a regulated carrier: change management through pull request and pipeline, segregation of duties between authoring and deploying, least-privilege access, and audit evidence that falls out of CI/CD rather than being assembled for an auditor. • Develop reusable frameworks, templates, and patterns that raise the pod's consistency and delivery speed. Delivery leadership • Partner with the Product Owner and Scrum Master on decomposition and refinement: turn use cases and features into estimable, grounded stories with testable acceptance criteria and an identified target layer and repository. • Hold the Definition of Ready before the pod commits and the Definition of Done before the pod calls something finished - merged and approved code, passing CI and coverage gates, and evidence that the outcome is real. • Identify unknowns that need a spike rather than an estimate, and say so during planning rather than mid-sprint. Mentorship and collaboration • Grow the engineers on the pod through design review, pairing, and code review rather than by taking the hard work yourself. • Bring new engineers up to productive speed on the platform's conventions and tooling, and reduce single points of knowledge - no data product that only one person understands. • Work with the platform team on capability gaps: when the pod needs something the platform does not yet offer, raise it as a demand signal rather than building a private workaround. • Partner with the DataOps/MLOps Lead on the shared CI/CD, orchestration, and observability standards - adopt and strengthen the paved path rather than forking it, and be the pod's voice on what it is still missing. • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements. Qualifications Required Qualifications • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered. • 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through to published data products. • Strong hands-on Python and SQL, with the credibility to make design calls and the willingness to still write and review code. • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning. • Deep dimensional modeling experience - grain, keys, slowly changing dimensions, facts and dimensions, conformed dimensions. • Demonstrated technical leadership: setting standards, leading design, and raising other engineers' work, whether or not the role carried a lead title. • Experience owning data that other teams depend on, including handling breaking changes and production data incidents. • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD - Azure DevOps or comparable. • Experience setting observability and SLA/SLO expectations for data other teams depend on, and running the incident and communication path when they are missed. • Ability to explain trade-offs clearly to engineers, product owners, and business stakeholders, and to say no to a design that will not hold. Preferred Qualifications • Databricks certification (Data Engineer Professional or equivalent demonstrated depth). • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage. • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake. • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments. • Practical experience with data contracts, ODCS, or data-mesh style data product ownership. • Experience with a declarative Python ingestion framework such as dlt (dltHub) or comparable. • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar. • Familiarity with MLOps practice - MLflow, model registries, and model serving - sufficient to design data products that ML systems can depend on. • Azure and Azure DevOps. • Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations. • Experience introducing AI-assisted development into a team's normal workflow in a disciplined way. $109,500 - $167,833 a year Protective's targeted salary range for this position is $109,500 to $167,833. Actual salaries may vary depending on factors, including but not limited to, job location, skills, and experience. The range listed is just one component of Protective's total compensation package for employees. This position also offers additional incentive opportunities through an annual incentive based on individual and Company performance. Employee Benefits: We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits ( e.g. , paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health. Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective's platform to improve wellbeing while earning cash rewards. Eligibility for certain benefits may vary by position in accordance with the terms of the Company's benefit plans. Accommodations for Applicants with a Disability: If you require an accommodation to complete the application and recruitment process due to a disability, please email eric . click apply for full job details
09/23/2026
Full time
The work we do has an impact on millions of lives, and you can be a part of it. We help protect our customers against life's uncertainties. Regardless of where you work within the company, you'll be helping provide protection and peace of mind when our customers need it most. Protective is looking for a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, our Databricks lakehouse on Azure. You will lead the design of data products through the full medallion architecture - Bronze ingestion, Silver conformance, and Gold consumption - and be accountable for whether those products hold up for the consumers who depend on them. This is a hands-on technical leadership role, not a people-management or project-management role. You will still write and review production code daily. What you own is how the pod's data products are designed, modeled, contracted, and tested, and the standard the pod holds itself to. The Product Owner owns the backlog and the Scrum Master owns the sprint; you own the engineering. On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description. Key Responsibilities Design and modeling • Lead the design of data products end to end: what gets ingested, how it is cleansed and conformed, how it is modeled, and what the Gold layer looks like to the people querying it. • Own dimensional design - grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across the pod's products. • Set the pod's position on where logic belongs: what is cleaned in Silver, what is business logic in Gold, and what is a consumer's own concern. • Keep models as simple as the questions require, and push back on designs that will not hold. • Partner with ML engineering where the pod's Gold layer is the training or feature source for a model, so those datasets are contracted, versioned, and reproducible like any other consumer-facing product. Data contracts and consumer compatibility • Own the pod's ODCS data contracts as real interfaces: named owners, named consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy. • Make the compatibility call on every proposed contract change, and drive consumer notification when a change is genuinely breaking. • Represent the pod's contracts in cross-domain conversations, where one pod's Gold layer is another team's dependency. Standards, quality, and operations • Set and hold the pod's engineering standards for Python, SQL, dbt, testing, model structure, naming, and repository conventions - consistent with the platform's paved paths and Azure DevOps CI gates rather than in competition with them. • Lead code review. Be the reviewer who catches the modeling mistake, the missing test, and the change that will break a consumer, and who explains why so the pod learns it. • Ensure quality rules are enforced in tests and asset checks rather than asserted in documentation, and that pipeline health is observable without someone going to look - freshness, volume, latency, and cost instrumented, and alerting set against the SLAs and SLOs the pod's contracts commit to. • Own the pod's operational posture for its own pipelines: failure diagnosis, data-issue triage, backfills, cost and performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks someone other than the author can execute. • Keep the pod's delivery inside the control expectations of a regulated carrier: change management through pull request and pipeline, segregation of duties between authoring and deploying, least-privilege access, and audit evidence that falls out of CI/CD rather than being assembled for an auditor. • Develop reusable frameworks, templates, and patterns that raise the pod's consistency and delivery speed. Delivery leadership • Partner with the Product Owner and Scrum Master on decomposition and refinement: turn use cases and features into estimable, grounded stories with testable acceptance criteria and an identified target layer and repository. • Hold the Definition of Ready before the pod commits and the Definition of Done before the pod calls something finished - merged and approved code, passing CI and coverage gates, and evidence that the outcome is real. • Identify unknowns that need a spike rather than an estimate, and say so during planning rather than mid-sprint. Mentorship and collaboration • Grow the engineers on the pod through design review, pairing, and code review rather than by taking the hard work yourself. • Bring new engineers up to productive speed on the platform's conventions and tooling, and reduce single points of knowledge - no data product that only one person understands. • Work with the platform team on capability gaps: when the pod needs something the platform does not yet offer, raise it as a demand signal rather than building a private workaround. • Partner with the DataOps/MLOps Lead on the shared CI/CD, orchestration, and observability standards - adopt and strengthen the paved path rather than forking it, and be the pod's voice on what it is still missing. • Work with data architecture and governance on solution shape, Unity Catalog placement, and access requirements. Qualifications Required Qualifications • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered. • 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through to published data products. • Strong hands-on Python and SQL, with the credibility to make design calls and the willingness to still write and review code. • Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning. • Deep dimensional modeling experience - grain, keys, slowly changing dimensions, facts and dimensions, conformed dimensions. • Demonstrated technical leadership: setting standards, leading design, and raising other engineers' work, whether or not the role carried a lead title. • Experience owning data that other teams depend on, including handling breaking changes and production data incidents. • Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD - Azure DevOps or comparable. • Experience setting observability and SLA/SLO expectations for data other teams depend on, and running the incident and communication path when they are missed. • Ability to explain trade-offs clearly to engineers, product owners, and business stakeholders, and to say no to a design that will not hold. Preferred Qualifications • Databricks certification (Data Engineer Professional or equivalent demonstrated depth). • Unity Catalog at multi-team scale: catalogs, schemas, external locations, permissions, and lineage. • dbt at scale on Databricks, and Python-based modeling frameworks over Delta Lake. • Dagster and Dagster Cloud, including assets, asset checks, and branch deployments. • Practical experience with data contracts, ODCS, or data-mesh style data product ownership. • Experience with a declarative Python ingestion framework such as dlt (dltHub) or comparable. • Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar. • Familiarity with MLOps practice - MLflow, model registries, and model serving - sufficient to design data products that ML systems can depend on. • Azure and Azure DevOps. • Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations. • Experience introducing AI-assisted development into a team's normal workflow in a disciplined way. $109,500 - $167,833 a year Protective's targeted salary range for this position is $109,500 to $167,833. Actual salaries may vary depending on factors, including but not limited to, job location, skills, and experience. The range listed is just one component of Protective's total compensation package for employees. This position also offers additional incentive opportunities through an annual incentive based on individual and Company performance. Employee Benefits: We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits ( e.g. , paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health. Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective's platform to improve wellbeing while earning cash rewards. Eligibility for certain benefits may vary by position in accordance with the terms of the Company's benefit plans. Accommodations for Applicants with a Disability: If you require an accommodation to complete the application and recruitment process due to a disability, please email eric . click apply for full job details
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM Anduril Intelligence Systems (AIS) is a lead provider of highly specialized engineering products for Intelligence Community (IC) customers. We work within the IC to understand their requirements and shape concepts of operation. We design, develop, and deliver exquisite capability across their mission set using commercially available and custom hardware and software. We provide critically needed capabilities that address our customers' most pressing national security requirements. ABOUT THE JOB We are looking for a Software Engineer (SWE) to join our rapidly growing team in Reston, Virginia. In this role you will be responsible for developing software across projects. This role may include desktop application development, porting and refactoring code bases, performance tuning, front-end and back-end development, and other software development and engineering tasks as needed. You will need to work within tight timelines and resource constraints. WHAT YOU'LL DO Write software for desktop and full-stack software products. Work at various levels of the software stack. Work with existing teams to maintain and update existing software systems. Actively contribute to the software development for critical tasks as needed to meet program deadlines. Adhere to software best practices and coding standards, perform code reviews, interact with revision control, build processes, and testing. Triage issues and investigate root cause failures. Collaborate closely with project lead and other technical leads across multiple projects to ensure software remains integrated and operational with other software development efforts. REQUIRED QUALIFICATIONS Proficiency in developing and maintaining C# code 3-5 years of experience in full-stack software development B.S. in Computer Science, Computer Engineering, or related fields. Experience working in constrained development environments, including air-gapped systems. Ability to quickly understand and navigate complex systems and established code bases. Ability to understand and implement Government certification requirements. Currently possesses and is able to maintain an active U.S. Top Secret security clearance with Full Scope Polygraph. PREFERRED QUALIFICATIONS Experience working with Windows Desktop application development frameworks such as WPF or WinForms Experience working in JavaScript (React). 5+ years of experience in full-stack software development. Experience working in sensitive, secure government spaces. US Salary Range $166,000-$220,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits. Protecting Yourself from Recruitment Scams Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. Please always verify communications: Direct from Anduril: If you receive an email from one of our recruiters, it will only come from address. Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to . Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain before providing any personal information or clicking on links. What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to . Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts. Data Privacy To view Anduril's candidate data privacy policy, please visit By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.
09/23/2026
Full time
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM Anduril Intelligence Systems (AIS) is a lead provider of highly specialized engineering products for Intelligence Community (IC) customers. We work within the IC to understand their requirements and shape concepts of operation. We design, develop, and deliver exquisite capability across their mission set using commercially available and custom hardware and software. We provide critically needed capabilities that address our customers' most pressing national security requirements. ABOUT THE JOB We are looking for a Software Engineer (SWE) to join our rapidly growing team in Reston, Virginia. In this role you will be responsible for developing software across projects. This role may include desktop application development, porting and refactoring code bases, performance tuning, front-end and back-end development, and other software development and engineering tasks as needed. You will need to work within tight timelines and resource constraints. WHAT YOU'LL DO Write software for desktop and full-stack software products. Work at various levels of the software stack. Work with existing teams to maintain and update existing software systems. Actively contribute to the software development for critical tasks as needed to meet program deadlines. Adhere to software best practices and coding standards, perform code reviews, interact with revision control, build processes, and testing. Triage issues and investigate root cause failures. Collaborate closely with project lead and other technical leads across multiple projects to ensure software remains integrated and operational with other software development efforts. REQUIRED QUALIFICATIONS Proficiency in developing and maintaining C# code 3-5 years of experience in full-stack software development B.S. in Computer Science, Computer Engineering, or related fields. Experience working in constrained development environments, including air-gapped systems. Ability to quickly understand and navigate complex systems and established code bases. Ability to understand and implement Government certification requirements. Currently possesses and is able to maintain an active U.S. Top Secret security clearance with Full Scope Polygraph. PREFERRED QUALIFICATIONS Experience working with Windows Desktop application development frameworks such as WPF or WinForms Experience working in JavaScript (React). 5+ years of experience in full-stack software development. Experience working in sensitive, secure government spaces. US Salary Range $166,000-$220,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits. Protecting Yourself from Recruitment Scams Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. Please always verify communications: Direct from Anduril: If you receive an email from one of our recruiters, it will only come from address. Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to . Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain before providing any personal information or clicking on links. What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to . Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts. Data Privacy To view Anduril's candidate data privacy policy, please visit By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
09/23/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
09/23/2026
Full time
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at . About the Role CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management. You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work. You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation. Key Responsibilities Production Support and Troubleshooting Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures. Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately. Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data. Perform and validate approved remediation, ensuring services and devices return to a healthy state. Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions. Participate in the team's on-call rotation after completing onboarding and a readiness review. Observability and Reliability Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems. Maintain and improve dashboards, alerts, and operational KPIs for team-owned services. Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements. Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent. Validate software, firmware, or process changes in test environments and support controlled production rollouts. Documentation and Collaboration Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation. Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence. Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution. Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes. Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain. Minimum Qualifications Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience. Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting. Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment. Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language. Scripting experience in Bash, or another shell scripting language. Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure. A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication. Experience participating in an on-call rotation or supporting time-sensitive production issues. Preferred Qualifications Experience supporting Kubernetes applications or distributed production services. Experience with incident response, escalation, root-cause analysis, or post-incident reviews. Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management. Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations. Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation. Experience collaborating with hardware or firmware vendors to investigate and validate fixes. Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Why CoreWeave? At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core Act Like an Owner Empower Employees Deliver Best-in-Class Client Experiences Achieve More Together We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! The base salary range for this role is $109,000 to $145,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). What We Offer The range we've posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account Tuition Reimbursement Ability to Participate in Employee Stock Purchase Program (ESPP) Mental Wellness Benefits through Spring Health Family-Forming support provided by Carrot Paid Parental Leave Flexible, full-service childcare support with Kinside 401(k) with a generous employer match Flexible PTO Catered lunch each day in our office and data center locations A casual work environment A work culture focused on innovative disruption California Applicants California Consumer Privacy Act Equal Opportunity & Accommodations CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA) . click apply for full job details
Johnson & Johnson Innovative Medicine
Santa Clara, California
At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit. Job Function: R&D Product Development Job Sub Function: R&D Mechanical Engineering Job Category: Scientific/Technology All Job Posting Locations: Santa Clara, California, United States of America Job Description: The Robotics and Digital solutions (RAD) group, part of the Johnson & Johnson family of companies, is recruiting for a Hardware Reliability Test Engineer for the Robotics and Digital R&D Team. This position is located in Santa Clara, CA. What We Do: At Johnson & Johnson MedTech, we are building the future of robotic surgery and healthcare through innovation and technical excellence. Our goal is to develop advanced robotic platforms that are precise, reliable, and accessible. Our team focuses on creating solutions that empower surgeons and improve patient outcomes worldwide. Who We Are: Within RAD R&D we are a team of clinical, electrical, firmware, mechanical, mechatronics, robotic controls, systems and software engineers who are passionate about improving patient care. The team includes a wide range of experience levels from junior engineers to industry experts. We follow an iterative, collaborative approach to product development, working across clinical, instrument, accessories, and system teams. We prioritize autonomy, develop a culture of learning, and value diversity of thought and background. Our environment is inclusive, supportive, and driven by a shared commitment to innovation and excellence in healthcare. Core Job Responsibilities: Translate product reliability requirements and specifications into detailed test protocols and pass/fail criteria. Develop and implement mechanical, thermal, electrical, and environmental loads/stresses (e.g., vibration, shock, thermal cycling, humidity, electrical stress) required by the protocols. Plan and execute Reliability Demonstration Testing, Characterization Testing, and Ongoing Reliability Testing to validate product robustness across development and production lifecycle. Design, build, and validate custom test fixtures and test rigs to enable repeatable, controlled testing. Develop and maintain test automation and scripting for test setup, execution, data logging, and analysis (e.g., Python, LabVIEW, or equivalent). Run hands-on testing, monitor test runs, and ensure data integrity and traceability. Analyze test data, produce summary reports, and present findings to engineering and program teams. Document and report issues discovered during testing; create reproducible problem descriptions, supporting data and test artifacts. Triage issues to the appropriate project teams or Failure Analysis (FA) groups and support root-cause investigations as required. Maintain and calibrate test equipment and laboratory documentation; ensure test safety and compliance with internal standards and external regulations. Collaborate cross-functionally with design, manufacturing, quality, and FA to close reliability issues and feed lessons learned back into the product development cycle. Required Knowledge/Skills, Education, And Experience: Bachelor's degree in Mechanical Engineering, Electrical Engineering, Physics, or related technical field (or equivalent experience). 2+ years of hands-on experience in reliability or test engineering roles (design and execution of reliability tests, fixture design, or test automation). Practical experience with environmental/mechanical test equipment such as vibration tables, shock rigs, thermal chambers, and electrical stress equipment. Experience designing and building test fixtures and custom test rigs. Proficiency in scripting or programming for test automation and data capture (examples: Python, LabVIEW, or similar). Strong data analysis skills; experience with statistical analysis and data visualization (Excel, Python, MATLAB, Minitab, etc.). Clear technical writing skills for protocols, test plans, and test reports. Strong problem-solving skills and attention to detail; ability to prioritize multiple test projects. Preferred Knowledge/Skills, Education, And Experience: Master's degree in relevant engineering field Experience with reliability methods such as HALT/HASS, accelerated life testing (ALT). Familiarity with design for reliability concepts (FMEA, reliability prediction methods). Experience working in regulated industries (medical devices, aerospace, automotive) or with cross-functional FA teams. Experience with test equipment control and data acquisition hardware (DAQs, PXI, NI hardware) Working Conditions: Laboratory/shop environment with regular hands-on work on test fixtures and equipment. May require moderate lifting and use of power tools when building fixtures. Occasional travel to supplier sites, test labs, or failure analysis facilities may be required. The anticipated base pay range for this position is $106,000.00 to $170,200.00 The Company maintains highly competitive, performance-based compensation programs. Under current guidelines, this position is eligible for an annual performance bonus in accordance with the terms of the applicable plan. The annual performance bonus is a cash bonus intended to provide an incentive to achieve annual targeted results by rewarding for individual and the corporation's performance over a calendar/performance year. Bonuses are awarded at the Company's discretion on an individual basis. Employees and/or eligible dependents may be eligible to participate in the following Company sponsored employee benefit programs: medical, dental, vision, life insurance, short- and long-term disability, business accident insurance, and group legal insurance. Employees may be eligible to participate in the Company's consolidated retirement plan (pension) and savings plan (401(k . This position is eligible to participate in the Company's long-term incentive program. Employees are eligible for the following time off benefits: Vacation - up to 120 hours per calendar year Sick time - up to 40 hours per calendar year Holiday pay, including Floating Holidays - up to 13 days per calendar year Work, Personal and Family Time - up to 40 hours per calendar year For additional general information on Company benefits, please go to: Johnson & Johnson is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, age, national origin, disability, protected veteran status or other characteristics protected by federal, state or local law. We actively seek qualified candidates who are protected veterans and individuals with disabilities as defined under VEVRAA and Section 503 of the Rehabilitation Act. Johnson & Johnson is committed to providing an interview process that is inclusive of our applicants' needs. If you are an individual with a disability and would like to request an accommodation, external applicants please contact us via . internal employees contact AskGS to be directed to your accommodation resource. Required Skills: Data Analysis, Fixtures Design, Mechanical Testing, Test Automation Preferred Skills: Accelerated Life Testing, Analytical Reasoning, Auto-CAD Design, Design Thinking, FEMA, Highly Accelerated Life Test (HALT), Mechanical Engineering, Problem Solving, Process Oriented, Product Reliability, Quality Control Testing, Reliability Engineering, Reliability Predictions, Reliability Testing, Research and Development The anticipated base pay range for this position is : $106,000.00 - $170,200.00 Additional Description for Pay Transparency:
09/23/2026
Full time
At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit. Job Function: R&D Product Development Job Sub Function: R&D Mechanical Engineering Job Category: Scientific/Technology All Job Posting Locations: Santa Clara, California, United States of America Job Description: The Robotics and Digital solutions (RAD) group, part of the Johnson & Johnson family of companies, is recruiting for a Hardware Reliability Test Engineer for the Robotics and Digital R&D Team. This position is located in Santa Clara, CA. What We Do: At Johnson & Johnson MedTech, we are building the future of robotic surgery and healthcare through innovation and technical excellence. Our goal is to develop advanced robotic platforms that are precise, reliable, and accessible. Our team focuses on creating solutions that empower surgeons and improve patient outcomes worldwide. Who We Are: Within RAD R&D we are a team of clinical, electrical, firmware, mechanical, mechatronics, robotic controls, systems and software engineers who are passionate about improving patient care. The team includes a wide range of experience levels from junior engineers to industry experts. We follow an iterative, collaborative approach to product development, working across clinical, instrument, accessories, and system teams. We prioritize autonomy, develop a culture of learning, and value diversity of thought and background. Our environment is inclusive, supportive, and driven by a shared commitment to innovation and excellence in healthcare. Core Job Responsibilities: Translate product reliability requirements and specifications into detailed test protocols and pass/fail criteria. Develop and implement mechanical, thermal, electrical, and environmental loads/stresses (e.g., vibration, shock, thermal cycling, humidity, electrical stress) required by the protocols. Plan and execute Reliability Demonstration Testing, Characterization Testing, and Ongoing Reliability Testing to validate product robustness across development and production lifecycle. Design, build, and validate custom test fixtures and test rigs to enable repeatable, controlled testing. Develop and maintain test automation and scripting for test setup, execution, data logging, and analysis (e.g., Python, LabVIEW, or equivalent). Run hands-on testing, monitor test runs, and ensure data integrity and traceability. Analyze test data, produce summary reports, and present findings to engineering and program teams. Document and report issues discovered during testing; create reproducible problem descriptions, supporting data and test artifacts. Triage issues to the appropriate project teams or Failure Analysis (FA) groups and support root-cause investigations as required. Maintain and calibrate test equipment and laboratory documentation; ensure test safety and compliance with internal standards and external regulations. Collaborate cross-functionally with design, manufacturing, quality, and FA to close reliability issues and feed lessons learned back into the product development cycle. Required Knowledge/Skills, Education, And Experience: Bachelor's degree in Mechanical Engineering, Electrical Engineering, Physics, or related technical field (or equivalent experience). 2+ years of hands-on experience in reliability or test engineering roles (design and execution of reliability tests, fixture design, or test automation). Practical experience with environmental/mechanical test equipment such as vibration tables, shock rigs, thermal chambers, and electrical stress equipment. Experience designing and building test fixtures and custom test rigs. Proficiency in scripting or programming for test automation and data capture (examples: Python, LabVIEW, or similar). Strong data analysis skills; experience with statistical analysis and data visualization (Excel, Python, MATLAB, Minitab, etc.). Clear technical writing skills for protocols, test plans, and test reports. Strong problem-solving skills and attention to detail; ability to prioritize multiple test projects. Preferred Knowledge/Skills, Education, And Experience: Master's degree in relevant engineering field Experience with reliability methods such as HALT/HASS, accelerated life testing (ALT). Familiarity with design for reliability concepts (FMEA, reliability prediction methods). Experience working in regulated industries (medical devices, aerospace, automotive) or with cross-functional FA teams. Experience with test equipment control and data acquisition hardware (DAQs, PXI, NI hardware) Working Conditions: Laboratory/shop environment with regular hands-on work on test fixtures and equipment. May require moderate lifting and use of power tools when building fixtures. Occasional travel to supplier sites, test labs, or failure analysis facilities may be required. The anticipated base pay range for this position is $106,000.00 to $170,200.00 The Company maintains highly competitive, performance-based compensation programs. Under current guidelines, this position is eligible for an annual performance bonus in accordance with the terms of the applicable plan. The annual performance bonus is a cash bonus intended to provide an incentive to achieve annual targeted results by rewarding for individual and the corporation's performance over a calendar/performance year. Bonuses are awarded at the Company's discretion on an individual basis. Employees and/or eligible dependents may be eligible to participate in the following Company sponsored employee benefit programs: medical, dental, vision, life insurance, short- and long-term disability, business accident insurance, and group legal insurance. Employees may be eligible to participate in the Company's consolidated retirement plan (pension) and savings plan (401(k . This position is eligible to participate in the Company's long-term incentive program. Employees are eligible for the following time off benefits: Vacation - up to 120 hours per calendar year Sick time - up to 40 hours per calendar year Holiday pay, including Floating Holidays - up to 13 days per calendar year Work, Personal and Family Time - up to 40 hours per calendar year For additional general information on Company benefits, please go to: Johnson & Johnson is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, age, national origin, disability, protected veteran status or other characteristics protected by federal, state or local law. We actively seek qualified candidates who are protected veterans and individuals with disabilities as defined under VEVRAA and Section 503 of the Rehabilitation Act. Johnson & Johnson is committed to providing an interview process that is inclusive of our applicants' needs. If you are an individual with a disability and would like to request an accommodation, external applicants please contact us via . internal employees contact AskGS to be directed to your accommodation resource. Required Skills: Data Analysis, Fixtures Design, Mechanical Testing, Test Automation Preferred Skills: Accelerated Life Testing, Analytical Reasoning, Auto-CAD Design, Design Thinking, FEMA, Highly Accelerated Life Test (HALT), Mechanical Engineering, Problem Solving, Process Oriented, Product Reliability, Quality Control Testing, Reliability Engineering, Reliability Predictions, Reliability Testing, Research and Development The anticipated base pay range for this position is : $106,000.00 - $170,200.00 Additional Description for Pay Transparency:
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM Anduril Intelligence Systems (AIS) is a lead provider of highly specialized engineering products for Intelligence Community (IC) customers. We work within the IC to understand their requirements and shape concepts of operation. We design, develop, and deliver exquisite capability across their mission set using commercially available and custom hardware and software. We provide critically needed capabilities that address our customers' most pressing national security requirements. ABOUT THE JOB We are looking for a Software Engineer (SWE) to join our rapidly growing team in Reston, Virginia. In this role you will be responsible for developing software across projects that may range from full-stack enterprise software to user-facing software on embedded and custom hardware devices. This role may include writing middleware, porting and refactoring code bases between languages, performance tuning, front-end and back-end development, and other software development and engineering tasks as needed. You will need to work within tight timelines and resource constraints. WHAT YOU'LL DO Work directly with project managers to write software for full-stack software and user-facing software on custom and embedded hardware. Work at various levels of the software stack, from databases to GUI frameworks. Work with existing teams to maintain and update existing software systems. Provide software designs, estimates, and schedules as needed to program and project management. Actively contribute to the software development for critical tasks as needed to meet program deadlines. Adhere to software best practices and coding standards, perform code reviews, interact with revision control, build processes, and testing. Triage issues and investigate root cause failures. Report to the overall software lead for the project. REQUIRED QUALIFICATIONS 1-3 years of experience in full-stack software development B.S. in Computer Science, Computer Engineering, or related fields. Experience working in constrained development environments, including air-gapped systems. Ability to quickly understand and navigate complex systems and established code bases. Ability to understand and implement complex certification requirements. Currently possesses and is able to maintain an active U.S. Top Secret security clearance with Full Scope Polygraph. PREFERRED QUALIFICATIONS Some supplementary experience in embedded systems development. Experience working in C#, C++, Python, and/or JavaScript. Experience working with AI/ML frameworks. Some experience in full-stack software development. Strong focus on security. US Salary Range $129,000-$171,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits. Protecting Yourself from Recruitment Scams Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. Please always verify communications: Direct from Anduril: If you receive an email from one of our recruiters, it will only come from address. Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to . Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain before providing any personal information or clicking on links. What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to . Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts. Data Privacy To view Anduril's candidate data privacy policy, please visit By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.
09/23/2026
Full time
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM Anduril Intelligence Systems (AIS) is a lead provider of highly specialized engineering products for Intelligence Community (IC) customers. We work within the IC to understand their requirements and shape concepts of operation. We design, develop, and deliver exquisite capability across their mission set using commercially available and custom hardware and software. We provide critically needed capabilities that address our customers' most pressing national security requirements. ABOUT THE JOB We are looking for a Software Engineer (SWE) to join our rapidly growing team in Reston, Virginia. In this role you will be responsible for developing software across projects that may range from full-stack enterprise software to user-facing software on embedded and custom hardware devices. This role may include writing middleware, porting and refactoring code bases between languages, performance tuning, front-end and back-end development, and other software development and engineering tasks as needed. You will need to work within tight timelines and resource constraints. WHAT YOU'LL DO Work directly with project managers to write software for full-stack software and user-facing software on custom and embedded hardware. Work at various levels of the software stack, from databases to GUI frameworks. Work with existing teams to maintain and update existing software systems. Provide software designs, estimates, and schedules as needed to program and project management. Actively contribute to the software development for critical tasks as needed to meet program deadlines. Adhere to software best practices and coding standards, perform code reviews, interact with revision control, build processes, and testing. Triage issues and investigate root cause failures. Report to the overall software lead for the project. REQUIRED QUALIFICATIONS 1-3 years of experience in full-stack software development B.S. in Computer Science, Computer Engineering, or related fields. Experience working in constrained development environments, including air-gapped systems. Ability to quickly understand and navigate complex systems and established code bases. Ability to understand and implement complex certification requirements. Currently possesses and is able to maintain an active U.S. Top Secret security clearance with Full Scope Polygraph. PREFERRED QUALIFICATIONS Some supplementary experience in embedded systems development. Experience working in C#, C++, Python, and/or JavaScript. Experience working with AI/ML frameworks. Some experience in full-stack software development. Strong focus on security. US Salary Range $129,000-$171,000 USD The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits. Protecting Yourself from Recruitment Scams Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. Please always verify communications: Direct from Anduril: If you receive an email from one of our recruiters, it will only come from address. Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to . Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain before providing any personal information or clicking on links. What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to . Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts. Data Privacy To view Anduril's candidate data privacy policy, please visit By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.
Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into enterprise grade code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work. We're backed by multiple tier 1 investors, and have proven success as founders of previous start-ups. Our Culture Who we are: Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500. How we work: We move Blitzy Fast: Time is both our company's and our clients' most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally. Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission. Passion for Invention: We're pushing the frontier of what's possible, requiring constant innovation and iteration. We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships. About the Role We are hiring a Principal Engineer to take full, hands-on ownership of Blitzy's most critical production-grade systems and to deliver high-leverage features that materially improve customer outcomes and engineering velocity. This is the most senior individual contributor role at the company today. This is not a Senior-plus role, an architecture-only role, or a promotion-track role. We are looking for someone who has already operated at Principal / Staff+ scope in a highly technical environment and expects to spend their time writing, reviewing, and shipping production code. This role is 100% hands-on. Leverage comes from system ownership, execution quality, and durable technical decisions - not people management or process. Responsibilities Own mission-critical production systems end-to-end, ensuring correctness, scalability, performance, reliability, and operational excellence. Design, build, and ship high-impact backend systems and features that improve product reliability, performance, and customer value. Architect scalable services and cloud infrastructure using technologies such as Python, REST, gRPC, Kubernetes, and Terraform. Identify and resolve complex technical bottlenecks that limit engineering quality, system performance, or organizational velocity. Build and operate LLM-powered systems and validation loops that evaluate correctness, consistency, durability, and production performance. Design and evolve data architectures incorporating relational, NoSQL, graph, and vector databases to support complex enterprise applications and semantic retrieval. Modernize and improve complex enterprise systems while balancing reliability, maintainability, scalability, and delivery speed. Set and uphold engineering quality standards through hands-on technical leadership, sound technical judgment, and ownership of long-term technical decisions. Qualifications Direct experience with Python as a primary programming language, backend frameworks, and microservices architectures. Expertise in REST and gRPC, with proficiency in Node.js and JavaScript. Proficiency in AWS, along with experience using at least one additional cloud platform such as GCP or Azure. Advanced knowledge of Kubernetes and Terraform in production environments. Experience operating highly available production systems, including monitoring, scalability, reliability, performance optimization, and operational tooling. Strong knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, MongoDB, Cassandra, or DynamoDB. Familiarity with graph databases such as Neo4j and vector databases or embedding infrastructure for semantic search and retrieval. Hands-on experience building and operating LLM-powered systems in production, including evaluation, validation, regression testing, tracing, and failure analysis. Working knowledge of LangSmith or comparable LLM observability and evaluation tools; familiarity with OpenAI, Anthropic, or similar model providers is a plus. Ability to contribute across the full stack, with a strong understanding of frontend architecture and the ability to debug, design, and ship across frontend, backend, infrastructure, and AI systems. Understanding of large-scale enterprise software systems, including architecture, integration, deployment, modernization, and long-term maintainability. Proven track record of operating at Staff+, Principal Engineer, or comparable scope, independently driving complex technical initiatives and delivering high-impact outcomes with minimal supervision. Salary Range: $300,000 - $500,000 + Bonus + Equity Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger. Base salary ranges are determined by country, role, level, experience, and skills. The range displayed on each job posting reflects Blitzy's good faith determination of the minimum and maximum targets for new hire salaries across all US locations. Individual pay is determined by related factors, including job skills, experience, and relevant education or training, which may impact a final offer. Your Talent Partner can share more about the specific salary range during the hiring process.
09/23/2026
Full time
Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into enterprise grade code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work. We're backed by multiple tier 1 investors, and have proven success as founders of previous start-ups. Our Culture Who we are: Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500. How we work: We move Blitzy Fast: Time is both our company's and our clients' most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally. Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission. Passion for Invention: We're pushing the frontier of what's possible, requiring constant innovation and iteration. We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships. About the Role We are hiring a Principal Engineer to take full, hands-on ownership of Blitzy's most critical production-grade systems and to deliver high-leverage features that materially improve customer outcomes and engineering velocity. This is the most senior individual contributor role at the company today. This is not a Senior-plus role, an architecture-only role, or a promotion-track role. We are looking for someone who has already operated at Principal / Staff+ scope in a highly technical environment and expects to spend their time writing, reviewing, and shipping production code. This role is 100% hands-on. Leverage comes from system ownership, execution quality, and durable technical decisions - not people management or process. Responsibilities Own mission-critical production systems end-to-end, ensuring correctness, scalability, performance, reliability, and operational excellence. Design, build, and ship high-impact backend systems and features that improve product reliability, performance, and customer value. Architect scalable services and cloud infrastructure using technologies such as Python, REST, gRPC, Kubernetes, and Terraform. Identify and resolve complex technical bottlenecks that limit engineering quality, system performance, or organizational velocity. Build and operate LLM-powered systems and validation loops that evaluate correctness, consistency, durability, and production performance. Design and evolve data architectures incorporating relational, NoSQL, graph, and vector databases to support complex enterprise applications and semantic retrieval. Modernize and improve complex enterprise systems while balancing reliability, maintainability, scalability, and delivery speed. Set and uphold engineering quality standards through hands-on technical leadership, sound technical judgment, and ownership of long-term technical decisions. Qualifications Direct experience with Python as a primary programming language, backend frameworks, and microservices architectures. Expertise in REST and gRPC, with proficiency in Node.js and JavaScript. Proficiency in AWS, along with experience using at least one additional cloud platform such as GCP or Azure. Advanced knowledge of Kubernetes and Terraform in production environments. Experience operating highly available production systems, including monitoring, scalability, reliability, performance optimization, and operational tooling. Strong knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, MongoDB, Cassandra, or DynamoDB. Familiarity with graph databases such as Neo4j and vector databases or embedding infrastructure for semantic search and retrieval. Hands-on experience building and operating LLM-powered systems in production, including evaluation, validation, regression testing, tracing, and failure analysis. Working knowledge of LangSmith or comparable LLM observability and evaluation tools; familiarity with OpenAI, Anthropic, or similar model providers is a plus. Ability to contribute across the full stack, with a strong understanding of frontend architecture and the ability to debug, design, and ship across frontend, backend, infrastructure, and AI systems. Understanding of large-scale enterprise software systems, including architecture, integration, deployment, modernization, and long-term maintainability. Proven track record of operating at Staff+, Principal Engineer, or comparable scope, independently driving complex technical initiatives and delivering high-impact outcomes with minimal supervision. Salary Range: $300,000 - $500,000 + Bonus + Equity Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger. Base salary ranges are determined by country, role, level, experience, and skills. The range displayed on each job posting reflects Blitzy's good faith determination of the minimum and maximum targets for new hire salaries across all US locations. Individual pay is determined by related factors, including job skills, experience, and relevant education or training, which may impact a final offer. Your Talent Partner can share more about the specific salary range during the hiring process.
About Mercor Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents. Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You'll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role: We're hiring our first IT Operations Lead to own the reliability, security, and scalability of our identity, endpoint, and IT infrastructure systems. You'll partner directly with the Head of IT to transform reactive support into strategic, automated systems that prevent problems before they occur - balancing speed, security, and an exceptional user experience for technical teams building the future of AI. You'll be based in our San Francisco or New York office , collaborating closely with Security, Engineering, and People Ops. In This Role, You Will: Handle day-to-day operations and incident response Triage and resolve Tier 2/3 tickets with a focus on automation and improving response and resolution times Manage and troubleshoot Apple devices at hyper-scale Trace root causes across complex, multi-system failures (Rippling Okta Kandji Google Workspace), identify patterns in recurring tickets, and propose automation or self-service solutions Analyze downstream impacts before making changes; map dependencies and blast radius for SSO, MDM, and access-control changes Build and maintain runbooks, troubleshooting guides, and knowledge base articles that elevate team capabilities Lead small projects that address operational pain points (BYOD policies, international provisioning, compliance evidence collection) Translate technical issues for non-technical stakeholders (People Ops, Finance, Legal) during incidents and changes Mentor future IT team members on troubleshooting methodology and systems thinking Participate in an on-call rotation and help establish sustainable escalation procedures as the team grows You May Be a Good Fit If You: 5-7 years in IT Operations or Technical Support roles, ideally supporting technical teams (SaaS, cloud, AI/ML environments) Systems thinker who naturally traces dependencies, considers second-order effects, and asks "why did this break?" not just "how do I fix it?" Strong incident management skills: triage, root-cause analysis, blameless postmortems, pattern recognition Understanding of compliance controls (access management, logging, change management) Clear communicator who can explain technical issues to both engineers and non-technical stakeholders and write excellent documentation Self-directed with a bias to action and strong judgment on when to escalate versus resolve independently Solves problems others gave up on through creative, systematic troubleshooting Thinks through second- and third-order effects before making changes Documents solutions that help everyone, not just yourself Builds trust through technical competence and calm, clear communication under pressure Strong Candidates May Also Have Experience With: Expert troubleshooting across the Apple ecosystem, including MDM (Kandji, Jamf, Intune) Advanced Google Workspace and Okta administration (SAML/OIDC, lifecycle automation, SCIM provisioning) Multi-cloud Support (AWS, Azure, GCP) Network troubleshooting (DNS, VPNs, VLANs) Scripting and automation (Python, Bash) and APIs for repetitive tasks and integrations, plus low-code automation tools (Okta Workflows, Zapier) What Makes This Role Unique You'll spend most of your day handling tickets and incidents , but you won't just fix problems, you'll ask why they exist and work to prevent them. If you see the same issue three times, you'll write the runbook, propose the automation, or flag the upstream fix. As we scale, you'll hire and mentor the next IT Operations team members, shaping how we support a fast-growing AI company. Benefits Bi-annual performance bonus structure Generous equity grant vested over 4 years Up to $15k Relocation bonus $10K housing bonus (if you live within 0.5 miles of our office) $1.5K monthly stipend for meals Free Equinox membership $200 monthly laundry reimbursement $200 monthly personal wellness reimbursement Health, Dental, Vision insurance
09/23/2026
Full time
About Mercor Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents. Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You'll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role: We're hiring our first IT Operations Lead to own the reliability, security, and scalability of our identity, endpoint, and IT infrastructure systems. You'll partner directly with the Head of IT to transform reactive support into strategic, automated systems that prevent problems before they occur - balancing speed, security, and an exceptional user experience for technical teams building the future of AI. You'll be based in our San Francisco or New York office , collaborating closely with Security, Engineering, and People Ops. In This Role, You Will: Handle day-to-day operations and incident response Triage and resolve Tier 2/3 tickets with a focus on automation and improving response and resolution times Manage and troubleshoot Apple devices at hyper-scale Trace root causes across complex, multi-system failures (Rippling Okta Kandji Google Workspace), identify patterns in recurring tickets, and propose automation or self-service solutions Analyze downstream impacts before making changes; map dependencies and blast radius for SSO, MDM, and access-control changes Build and maintain runbooks, troubleshooting guides, and knowledge base articles that elevate team capabilities Lead small projects that address operational pain points (BYOD policies, international provisioning, compliance evidence collection) Translate technical issues for non-technical stakeholders (People Ops, Finance, Legal) during incidents and changes Mentor future IT team members on troubleshooting methodology and systems thinking Participate in an on-call rotation and help establish sustainable escalation procedures as the team grows You May Be a Good Fit If You: 5-7 years in IT Operations or Technical Support roles, ideally supporting technical teams (SaaS, cloud, AI/ML environments) Systems thinker who naturally traces dependencies, considers second-order effects, and asks "why did this break?" not just "how do I fix it?" Strong incident management skills: triage, root-cause analysis, blameless postmortems, pattern recognition Understanding of compliance controls (access management, logging, change management) Clear communicator who can explain technical issues to both engineers and non-technical stakeholders and write excellent documentation Self-directed with a bias to action and strong judgment on when to escalate versus resolve independently Solves problems others gave up on through creative, systematic troubleshooting Thinks through second- and third-order effects before making changes Documents solutions that help everyone, not just yourself Builds trust through technical competence and calm, clear communication under pressure Strong Candidates May Also Have Experience With: Expert troubleshooting across the Apple ecosystem, including MDM (Kandji, Jamf, Intune) Advanced Google Workspace and Okta administration (SAML/OIDC, lifecycle automation, SCIM provisioning) Multi-cloud Support (AWS, Azure, GCP) Network troubleshooting (DNS, VPNs, VLANs) Scripting and automation (Python, Bash) and APIs for repetitive tasks and integrations, plus low-code automation tools (Okta Workflows, Zapier) What Makes This Role Unique You'll spend most of your day handling tickets and incidents , but you won't just fix problems, you'll ask why they exist and work to prevent them. If you see the same issue three times, you'll write the runbook, propose the automation, or flag the upstream fix. As we scale, you'll hire and mentor the next IT Operations team members, shaping how we support a fast-growing AI company. Benefits Bi-annual performance bonus structure Generous equity grant vested over 4 years Up to $15k Relocation bonus $10K housing bonus (if you live within 0.5 miles of our office) $1.5K monthly stipend for meals Free Equinox membership $200 monthly laundry reimbursement $200 monthly personal wellness reimbursement Health, Dental, Vision insurance
Job Description Summary: Digital products play a central role in how we create value for customers, support the teams who serve them, and shape the consumer experience. Our product organization brings together small, empowered teams that move with clarity, speed, and purpose, enabling digital to be a meaningful source of advantage across Coca-Cola's North America Operating Unit. Our work spans customer journeys, service delivery, sales workflows, and the platforms that connect them. We are raising our standards for product craft and rebuilding the systems behind these experiences. As a Tech Lead specializing in Machine Learning and Data Engineering, you will lead the technical direction for end-to-end ML capabilities that ship as part of our product, while also ensuring the data foundations (events, pipelines, feature tables, and governance) are reliable and scalable. You'll partner with Product, Design, Data Science/Analytics, and platform teams to frame problems, define success metrics, and guide solutions from data modeling and feature engineering through model training, deployment, monitoring, and iteration. This is a hands-on leadership role for engineers who can set standards, unblock teams, and drive execution across the ML and data stack without formal people-management responsibilities. What You Will Work On: Build ML-powered data products that model transaction drivers and surface optimized actions as insights to be embedded within integrated internal and external digital experiences that shape how our beverage brands activate across retail, foodservice, and digital channels. The success of our products is tied directly to measurable transaction lift at the point of sale, a primary objective of the North America Operating Unit and The Coca-Cola Company as a whole. How We Work You'll be part of a dedicated, cross-functional team (Product, Design, Engineering) that is: Empowered to solve problems, not just build features Accountable for outcomes, not output Collaborative by default, from discovery through delivery Continuously learning, using data and customer insight to improve Key Responsibilities Technical direction for a product ML domain: problem framing, approach selection, evaluation strategy, and iteration Data and feature foundations: event/telemetry definitions, transformation logic, feature/label tables, and training/serving consistency Production ML systems: deployment patterns (batch/online), model performance/latency tradeoffs, and operational readiness Quality and reliability: data quality checks, model monitoring (drift/performance), alerting, and runbooks Engineering standards: design reviews, code review quality, documentation, and reusable patterns for ML + data workflows Mentorship and enablement: coaching engineers through complex work and unblocking delivery across teams Develop, Train & Evaluate Models Build baselines and iterate on model approaches appropriate to the product problem (e.g., gradient boosting, deep learning, ranking) Lead feature engineering with strong data discipline: define entities and joins, validate labels, and ensure training/serving consistency Run experiments and evaluate models using sound methodology (train/validation splits, cross-validation as appropriate, error analysis) Document findings and recommendations clearly for technical and non-technical audiences Deploy & Operate Models in Production Deploy models to production (batch and/or real-time) with attention to latency, reliability, and cost Implement monitoring for upstream data and feature freshness/quality, drift, and model performance; define alerting and response playbooks Automate repeatable training and evaluation workflows (versioning, reproducibility, and artifact tracking) Participate in incident response and post-incident reviews when model behavior impacts customers or operations Establish reusable patterns for feature pipelines (batch/stream), backfills, and schema evolution; raise the bar through design reviews Define and reinforce standards for data governance and responsible ML (PII handling, access controls, data contracts, bias/fairness considerations) Partner with platform teams on the data stack (warehouse/lakehouse, streaming, orchestration) and MLOps tooling (feature stores, training infrastructure, deployment, monitoring) What We're Looking For Applied ML fundamentals: Understands supervised learning, evaluation metrics, and common failure modes Strong programming skills: Comfortable in Python and writing production-quality code (testing, readability, performance) Data intuition: Able to analyze datasets with SQL and/or Python, spot issues, and reason about bias/leakage Product mindset: Cares about measurable impact, guardrails, and user experience-not just model metrics Cross-functional collaboration: Partners with Product, Data Science, and Engineering to ship and iterate on ML features MLOps + data platform fluency: Comfortable with deployment, monitoring, reproducibility, and the pipelines/warehouses/streams that feed models Key Qualifications 6+ years of experience in machine learning engineering, data engineering, or software engineering, including leading technical direction for ML/data systems Demonstrated ownership of model development and evaluation, including metric selection, error analysis, and experimentation discipline Strong engineering fundamentals in Python (and SQL) with production practices (testing, reviews, CI/CD); familiarity with ML frameworks (e.g., PyTorch/TensorFlow) and data tooling (e.g., Spark, dbt, Airflow/Dagster) is preferred Experience shipping and operating ML systems in production, including model monitoring, rollback/retraining strategies, and coordination with upstream data/feature pipelines Familiarity with data platforms (data warehouse/lakehouse concepts), and exposure to orchestration/ETL tools (e.g., Microsoft fabric, Airflow, dbt, Spark) Preferred Qualifications Experience building product ML systems such as personalization, recommendations, ranking, forecasting, or NLP Experience with experimentation and measurement (A/B testing, uplift/impact analysis, online guardrails) Experience with feature pipelines or feature stores, and patterns for training/serving consistency Experience designing and operating data pipelines that power ML (batch and streaming), with clear SLAs for freshness and quality Experience with lakehouse/warehouse modeling for analytics and ML (dimensional/event models, backfills, schema evolution, data contracts) Demonstrated tech lead behaviors: driving design reviews, setting standards, mentoring engineers, and aligning stakeholders on tradeoffs Experience with model and data observability (drift detection, performance monitoring, dashboards/alerting) Familiarity with responsible AI and data privacy considerations (PII handling, access controls, model risk) Experience with production infrastructure (e.g., Docker/Kubernetes) or workflow tooling (e.g., Airflow, Dagster) used to run ML jobs Familiarity with modern engineering practices (CI/CD, testing, observability) Education Bachelor's degree in Computer Science, Engineering, or a related field Equivalent practical experience is equally valued Who Thrives Here Enjoy leading through influence-turning ambiguous problems into clear ML + data plans and helping others execute Communicate clearly across Product, Data Science, Analytics, and Engineering-especially around definitions, tradeoffs, and risk Take pride in raising the bar: reliable models and data pipelines, strong documentation, and operational follow-through Who This Role Is Not For This role may not be the right fit if you: Want to focus only on research prototypes or only on data pipelines (instead of owning end-to-end product ML systems) Avoid leading through influence (design reviews, alignment, mentorship) and prefer not to set or uphold technical standards Prefer to avoid operational responsibility for model and data health (monitoring, incidents, data quality/freshness, and continuous improvement) The Coca-Cola Company will not offer sponsorship for employment status (including, but not limited to, H1-B visa status and other employment-based nonimmigrant visas) for this position. Accordingly, all applicants must be currently authorized to work in the United States on a full-time basis and must not require The Coca-Cola Company's sponsorship to continue to work legally in the United States. Skills: Agile Methodology, Atlassian JIRA, Business Processes, Business Process Modeling, Cloud Platform, Communication, Data Flow Diagram, DevOps, Digital Transformation, Enterprise Architecture Framework, Enterprise Content Management (ECM), Java (Programming Language) . click apply for full job details
09/23/2026
Full time
Job Description Summary: Digital products play a central role in how we create value for customers, support the teams who serve them, and shape the consumer experience. Our product organization brings together small, empowered teams that move with clarity, speed, and purpose, enabling digital to be a meaningful source of advantage across Coca-Cola's North America Operating Unit. Our work spans customer journeys, service delivery, sales workflows, and the platforms that connect them. We are raising our standards for product craft and rebuilding the systems behind these experiences. As a Tech Lead specializing in Machine Learning and Data Engineering, you will lead the technical direction for end-to-end ML capabilities that ship as part of our product, while also ensuring the data foundations (events, pipelines, feature tables, and governance) are reliable and scalable. You'll partner with Product, Design, Data Science/Analytics, and platform teams to frame problems, define success metrics, and guide solutions from data modeling and feature engineering through model training, deployment, monitoring, and iteration. This is a hands-on leadership role for engineers who can set standards, unblock teams, and drive execution across the ML and data stack without formal people-management responsibilities. What You Will Work On: Build ML-powered data products that model transaction drivers and surface optimized actions as insights to be embedded within integrated internal and external digital experiences that shape how our beverage brands activate across retail, foodservice, and digital channels. The success of our products is tied directly to measurable transaction lift at the point of sale, a primary objective of the North America Operating Unit and The Coca-Cola Company as a whole. How We Work You'll be part of a dedicated, cross-functional team (Product, Design, Engineering) that is: Empowered to solve problems, not just build features Accountable for outcomes, not output Collaborative by default, from discovery through delivery Continuously learning, using data and customer insight to improve Key Responsibilities Technical direction for a product ML domain: problem framing, approach selection, evaluation strategy, and iteration Data and feature foundations: event/telemetry definitions, transformation logic, feature/label tables, and training/serving consistency Production ML systems: deployment patterns (batch/online), model performance/latency tradeoffs, and operational readiness Quality and reliability: data quality checks, model monitoring (drift/performance), alerting, and runbooks Engineering standards: design reviews, code review quality, documentation, and reusable patterns for ML + data workflows Mentorship and enablement: coaching engineers through complex work and unblocking delivery across teams Develop, Train & Evaluate Models Build baselines and iterate on model approaches appropriate to the product problem (e.g., gradient boosting, deep learning, ranking) Lead feature engineering with strong data discipline: define entities and joins, validate labels, and ensure training/serving consistency Run experiments and evaluate models using sound methodology (train/validation splits, cross-validation as appropriate, error analysis) Document findings and recommendations clearly for technical and non-technical audiences Deploy & Operate Models in Production Deploy models to production (batch and/or real-time) with attention to latency, reliability, and cost Implement monitoring for upstream data and feature freshness/quality, drift, and model performance; define alerting and response playbooks Automate repeatable training and evaluation workflows (versioning, reproducibility, and artifact tracking) Participate in incident response and post-incident reviews when model behavior impacts customers or operations Establish reusable patterns for feature pipelines (batch/stream), backfills, and schema evolution; raise the bar through design reviews Define and reinforce standards for data governance and responsible ML (PII handling, access controls, data contracts, bias/fairness considerations) Partner with platform teams on the data stack (warehouse/lakehouse, streaming, orchestration) and MLOps tooling (feature stores, training infrastructure, deployment, monitoring) What We're Looking For Applied ML fundamentals: Understands supervised learning, evaluation metrics, and common failure modes Strong programming skills: Comfortable in Python and writing production-quality code (testing, readability, performance) Data intuition: Able to analyze datasets with SQL and/or Python, spot issues, and reason about bias/leakage Product mindset: Cares about measurable impact, guardrails, and user experience-not just model metrics Cross-functional collaboration: Partners with Product, Data Science, and Engineering to ship and iterate on ML features MLOps + data platform fluency: Comfortable with deployment, monitoring, reproducibility, and the pipelines/warehouses/streams that feed models Key Qualifications 6+ years of experience in machine learning engineering, data engineering, or software engineering, including leading technical direction for ML/data systems Demonstrated ownership of model development and evaluation, including metric selection, error analysis, and experimentation discipline Strong engineering fundamentals in Python (and SQL) with production practices (testing, reviews, CI/CD); familiarity with ML frameworks (e.g., PyTorch/TensorFlow) and data tooling (e.g., Spark, dbt, Airflow/Dagster) is preferred Experience shipping and operating ML systems in production, including model monitoring, rollback/retraining strategies, and coordination with upstream data/feature pipelines Familiarity with data platforms (data warehouse/lakehouse concepts), and exposure to orchestration/ETL tools (e.g., Microsoft fabric, Airflow, dbt, Spark) Preferred Qualifications Experience building product ML systems such as personalization, recommendations, ranking, forecasting, or NLP Experience with experimentation and measurement (A/B testing, uplift/impact analysis, online guardrails) Experience with feature pipelines or feature stores, and patterns for training/serving consistency Experience designing and operating data pipelines that power ML (batch and streaming), with clear SLAs for freshness and quality Experience with lakehouse/warehouse modeling for analytics and ML (dimensional/event models, backfills, schema evolution, data contracts) Demonstrated tech lead behaviors: driving design reviews, setting standards, mentoring engineers, and aligning stakeholders on tradeoffs Experience with model and data observability (drift detection, performance monitoring, dashboards/alerting) Familiarity with responsible AI and data privacy considerations (PII handling, access controls, model risk) Experience with production infrastructure (e.g., Docker/Kubernetes) or workflow tooling (e.g., Airflow, Dagster) used to run ML jobs Familiarity with modern engineering practices (CI/CD, testing, observability) Education Bachelor's degree in Computer Science, Engineering, or a related field Equivalent practical experience is equally valued Who Thrives Here Enjoy leading through influence-turning ambiguous problems into clear ML + data plans and helping others execute Communicate clearly across Product, Data Science, Analytics, and Engineering-especially around definitions, tradeoffs, and risk Take pride in raising the bar: reliable models and data pipelines, strong documentation, and operational follow-through Who This Role Is Not For This role may not be the right fit if you: Want to focus only on research prototypes or only on data pipelines (instead of owning end-to-end product ML systems) Avoid leading through influence (design reviews, alignment, mentorship) and prefer not to set or uphold technical standards Prefer to avoid operational responsibility for model and data health (monitoring, incidents, data quality/freshness, and continuous improvement) The Coca-Cola Company will not offer sponsorship for employment status (including, but not limited to, H1-B visa status and other employment-based nonimmigrant visas) for this position. Accordingly, all applicants must be currently authorized to work in the United States on a full-time basis and must not require The Coca-Cola Company's sponsorship to continue to work legally in the United States. Skills: Agile Methodology, Atlassian JIRA, Business Processes, Business Process Modeling, Cloud Platform, Communication, Data Flow Diagram, DevOps, Digital Transformation, Enterprise Architecture Framework, Enterprise Content Management (ECM), Java (Programming Language) . click apply for full job details
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Compute Team is tasked with a critical and exciting mission: We deliver the compute platform responsible for running the fully autonomous vehicle's software stack. To achieve our mission, we architect and create high-performance custom silicon; we develop system-level compute architectures that push the boundaries of performance, power, and latency; and we collaborate closely with many other teammates to ensure we design and optimize hardware and software for maximum performance. We are a multidisciplinary team seeking curious and talented teammates to work on one of the world's highest performance automotive compute platforms. This role follows a hybrid work schedule and you will report to a Silicon Engineering Lead. You will: Collaborate with the Design, Verification, and Software teams to simulate future silicon designs and software on an emulation platform, targeting functional and performance validation and left-shift of software development Design, implement, and optimize emulation testbenches, balancing performance and debug capabilities based on user needs and hardware constraints Write end-to-end synthesizable transactors to interact with emulated designs through software APIs and C-DPI Triage and root-cause failures alongside emulation users, leveraging waveforms, software logging, and custom-built emulation monitors Assist with post-silicon bring-up, debug, and characterization Build infrastructure to support the emulation user base, including tools for model build, regression, automation, continuous integration, data analysis, and training You have: BS degree in Computer Science / Electrical Engineering or related field and 5 years of silicon development experience. Experience with at least one major hardware emulation platform (Palladium, Zebu, Veloce, Protium, HAPS). High level of proficiency with one or more of the following emulation approaches: virtual prototyping, testbench acceleration, hybrid emulation, in-circuit emulation (speed bridge), QEMU, or VirtualBox Experience with SystemVerilog and design verification methodologies Programming and scripting (C++, Python, or TCL) for automation, test development, debug flows, and release process. Strong debug and problem-solving ability across hardware and software We prefer: Experience with AXI/AMBA, PCIe, DRAM, and Ethernet interfaces Performance and power analysis techniques Experience with JTAG, DFT, UART, SPI, GPIO and other test/low-speed interfaces Knowledge of advanced design verification methods (coverage, gate-level simulation, assertions, and UVM). Post-silicon debug software (e.g. Trace32, OpenOCD, TARMAC) and lab bench tools (analyzers, scopes, meters) Understanding of bare metal programming, embedded systems, Linux internals, operating systems, boot loaders, drivers, and firmware The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $175,000-$215,000 USD
09/23/2026
Full time
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver -to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo's fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Compute Team is tasked with a critical and exciting mission: We deliver the compute platform responsible for running the fully autonomous vehicle's software stack. To achieve our mission, we architect and create high-performance custom silicon; we develop system-level compute architectures that push the boundaries of performance, power, and latency; and we collaborate closely with many other teammates to ensure we design and optimize hardware and software for maximum performance. We are a multidisciplinary team seeking curious and talented teammates to work on one of the world's highest performance automotive compute platforms. This role follows a hybrid work schedule and you will report to a Silicon Engineering Lead. You will: Collaborate with the Design, Verification, and Software teams to simulate future silicon designs and software on an emulation platform, targeting functional and performance validation and left-shift of software development Design, implement, and optimize emulation testbenches, balancing performance and debug capabilities based on user needs and hardware constraints Write end-to-end synthesizable transactors to interact with emulated designs through software APIs and C-DPI Triage and root-cause failures alongside emulation users, leveraging waveforms, software logging, and custom-built emulation monitors Assist with post-silicon bring-up, debug, and characterization Build infrastructure to support the emulation user base, including tools for model build, regression, automation, continuous integration, data analysis, and training You have: BS degree in Computer Science / Electrical Engineering or related field and 5 years of silicon development experience. Experience with at least one major hardware emulation platform (Palladium, Zebu, Veloce, Protium, HAPS). High level of proficiency with one or more of the following emulation approaches: virtual prototyping, testbench acceleration, hybrid emulation, in-circuit emulation (speed bridge), QEMU, or VirtualBox Experience with SystemVerilog and design verification methodologies Programming and scripting (C++, Python, or TCL) for automation, test development, debug flows, and release process. Strong debug and problem-solving ability across hardware and software We prefer: Experience with AXI/AMBA, PCIe, DRAM, and Ethernet interfaces Performance and power analysis techniques Experience with JTAG, DFT, UART, SPI, GPIO and other test/low-speed interfaces Knowledge of advanced design verification methods (coverage, gate-level simulation, assertions, and UVM). Post-silicon debug software (e.g. Trace32, OpenOCD, TARMAC) and lab bench tools (analyzers, scopes, meters) Understanding of bare metal programming, embedded systems, Linux internals, operating systems, boot loaders, drivers, and firmware The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process. Waymo employees are also eligible to participate in Waymo's discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $175,000-$215,000 USD
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details
09/23/2026
Full time
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack - energy, data centres, GPU superclusters, orchestration, and AI services - delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. About the Role (Job Purpose) Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands-on L2/L3 role operating at the intersection of GPU hardware, east-west networking, Linux, and data centre operations - acting as the operational bridge between Support, DC Operations, and Engineering. You will: Own complex, ambiguous problems end-to-end and make decisive calls in a results-driven environment, taking calculated risks where speed matters. Communicate technical detail clearly, specifically, and concisely - to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill. Influence without authority and build strong relationships with senior stakeholders across the business to get things done. Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast. Bring discipline and organisation: evidence-led investigations, accurate records, clean handovers. Experience required: 6+ years in infrastructure, operations, or support engineering roles in production environments, including 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates. What You'll be Doing (Responsibilities) Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes. Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA. Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing. Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion. Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans. Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation. Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover. Design and implement automation scripts and small tools to reduce toil and human intervention. Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter. Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews. Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion. Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise. About You (Skills / Qualifications Experience) Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2-3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity. Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch. GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA. High-performance east-west fabrics. Hands-on experience with RDMA fabrics - InfiniBand and/or RoCE - including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters. HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures. Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production. Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets. Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south. Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews. Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans. SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools. Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar). Data centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling - ideally gained through an operational support background. Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve. Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work. Nice to Have High-performance storage. Hands-on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage-network interaction and data-path performance issues. OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation, or similar). Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role. Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production. Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus. What We Can Offer You At Nscale, you'll find a collaborative, supportive . click apply for full job details