Job Description
Duration: 12-Month Contract + Possible Extension Industry: Energy / Utilities Seeking an experienced Geospatial Data Engineer with strong expertise in AWS, PySpark, and large-scale geospatial data processing. The ideal candidate will have hands-on experience building cloud-native data pipelines that support geospatial analytics, remote sensing initiatives, asset management, and risk modeling programs. Position Summary The Geospatial Data Engineer will partner with cross-functional teams including Data Engineering, Data Science, GIS, and Solution Architecture to design, develop, and optimize scalable geospatial data solutions. This role focuses heavily on AWS-based data engineering, large-scale raster and satellite imagery processing, and the development of high-performance geospatial analytics pipelines. The successful candidate will play a key role in building and maintaining data platforms that integrate asset, environmental, operational, and geospatial datasets to support advanced analytics and business intelligence initiatives. Core Responsibilities Data Pipeline Engineering Design, develop, and maintain scalable AWS-native data pipelines using Python and PySpark. Build automated ingestion, transformation, and processing workflows for large geospatial and operational datasets. Optimize pipeline performance, scalability, and reliability. Remote Sensing & Raster Data Processing Develop and support data pipelines for large raster-based datasets and satellite imagery. Manage multi-band imagery, including visible, near-infrared, and red-edge spectral bands. Implement selective and incremental ingestion strategies to efficiently process only relevant data subsets. Cloud Data Architecture Design and maintain cloud-native geospatial data solutions within AWS. Support data lake, warehouse, and analytical platform initiatives. Geospatial Processing Develop distributed geospatial processing solutions using Apache Sedona, GeoPandas, Shapely, and related technologies. Perform large-scale spatial analysis, joins, indexing, and optimization. Data Platform Development Support modern data engineering practices including CI/CD, automated testing, version control, and Infrastructure as Code. Contribute to enterprise data governance and data quality initiatives. Agile Collaboration Work closely with Business Analysts, Product Owners, GIS Specialists, Data Scientists, and Engineering teams in an Agile environment. Participate in sprint planning, design reviews, and technical discussions. Required Qualifications Education Bachelor's degree in Computer Science, Engineering, GIS, Geography, Data Science, or a related field. Experience 7+ years of Data Engineering experience designing and supporting enterprise-scale data pipelines. Technical Requirements Strong proficiency with Python, PySpark, SQL, and Apache Sedona. Hands-on experience building geospatial and raster data processing solutions on AWS. Experience developing cloud-native ETL and data integration workflows. Strong understanding of coordinate reference systems, projections, and spatial transformations (WGS84, NAD83, EPSG standards). Geospatial Technologies Experience with Shapefile, GeoJSON, GeoParquet, GeoPackage, KML, GeoTIFF, and Cloud-Optimized GeoTIFF (COG). Experience processing large-scale raster datasets and satellite imagery. Expertise with raster-vector analysis and multi-band imagery processing. Understanding of vegetation, environmental, and remote sensing analytics workflows. Spatial Analytics Spatial indexing and partitioning techniques including R-Tree, QuadTree, and distributed spatial joins. Geometry operations including buffering, intersections, nearest-neighbor analysis, topology validation, and geometry simplification. Experience addressing performance optimization challenges for large spatial datasets. Data Engineering & Orchestration Experience with Airflow, Dagster orchestration platforms. Knowledge of dimensional modeling, historical data management, and data warehousing concepts. Familiarity with CI/CD practices, automated testing, and Git-based development workflows. Nice to Have Experience with Palantir Foundry. STAC or other satellite imagery cataloging standards. LiDAR datasets and processing workflows. Utility, energy, environmental, infrastructure, or asset management industry experience. Experience building geospatial machine learning or advanced analytics solutions.