A

Sr. Data Platform Engineer

Aarden AI
1 day ago
Full-time
On-site
Seattle, Washington, United States
$150,000 - $180,000 USD yearly

About Us

Aarden is a land intelligence platform that helps landowners, investors, and developers figure out what a piece of land can actually be used for, and how to market it. We turn messy parcel, infrastructure, market, community, and ecological data into clear, bankable answers for land-dependent assets. Our goal is to become the default decision layer for land: helping physical projects start in places where they can be built and supported for decades.

Weโ€™ve built out a suite of data products to support that goal โ€” pipelines, databases, and AI/ML models that power our maps and In-app agents. We have strong product-market fit, and we're now focused on augmenting our data systems. Thatโ€™s where you come in.

The role

We're looking for a Product-Focused data engineer to maintain and evolve our geospatial data pipeline. Our core data asset is a unique blend of property data, geospatial data, and AI-native derived data. Alongside advocating for data excellence, youโ€™ll be empowered to opportunistically contribute to our user-facing product.

What youโ€™ll do

Pipeline modernization
  • Continue our migration of pipeline orchestration to Prefect
  • Own day-to-day operations of our data infrastructure
  • Extend the pipeline to include new data sources and transformations
  • Maintain, expand and optimize our postgres database and Iceberg datalake
  • Create the 'connective tissue' for data at Aarden
Cross-team integration
  • Partner with product on new feature-driven datasets.
  • Collaborate with the ML/analytics team to close the loop: anomaly detection โ†’ ticket โ†’ fix โ†’ validation โ†’ promotion to production
  • Develop cross-team tooling/infra to keep GitHub, Notion, Linear, and Slack connected so pipeline issues, docs, and fixes stay linked
Observability & AI-agent readiness
  • Implement run-over-run data observability (row counts, key column distributions) to catch anomalies and bugs
  • Expose accuracy/quality metrics as first-class artifacts so changes can be evaluated automatically, by a human or an agent
  • Write and maintain AI-context documentation (schema docs, pipeline architecture, known patterns/quirks, "what not to do")


You might be a good fit if youโ€ฆ

Must-have
  • Have strong Python skills & are comfortable with PySpark or similar distributed data processing
  • Have a strong sense of how the data youโ€™re working with impacts the end-user
  • Are curious and excited about AI and the impact it can have on our ways of working as developers
  • Have experience with geospatial data (GeoParquet, PostGIS, Apache Sedona, or similar)
  • Have worked with table formats like Apache Iceberg and lakehouse architectures
  • Have worked on workflow orchestration (Prefect, Airflow, Dagster, or similar)
  • Are comfortable working in a git-based, CI-friendly workflow
Strongly preferred
  • Have worked in full-stack environments, where your work can directly impact the application layer
  • Have experience with Apache Sedona or other cloud spatial-compute platforms
  • Have built observability/logging layers for data pipelines (not just app services)
  • Have experience with property, parcel, real estate, or land data specifically
Nice To have
  • Have experience using AI agents to improve data architecture in a real production codebase
  • Experience in real estate/land, energy, forestry, or agriculture tech


Our Stack

Languages: Python and SQL. TypeScript/Node is a plus for our Application layer and AWS ingest paths.
Orchestration & compute
  • Prefect 3 for pipeline orchestration (YAML/config-driven flows, retries, logging)
  • Coiled for elastic EC2 workers on GDAL-heavy and batch Python jobs
  • Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work
Data lake & formats: Apache Iceberg, Cloud-Optimized GeoTIFF (COG), and PMTiles. Queried with PySpark and DuckDB.
Geospatial: GDAL, rasterio, GeoPandas, and tippecanoe. Large-scale spatial work runs on Sedona/Spark via Wherobots.
Databases & serving: PostgreSQL + PostGIS (and pgvector on the app side) as the production store.


Working at Aarden

Aarden is a high-trust, high-output team. Weโ€™re striving to be intentional about our team growth. This allows us to test the outer boundaries of our individual capabilities, while also going deeper on developer tooling and support. Youโ€™ll work hard here, and weโ€™ve got your back.
Practically, this means youโ€™ll be asked to take on large projects, have a high bar of expectations to meet, and have a strong support system to help you meet that high bar. That support system includes:
  • At least 2 in-person days per week at our office in Capitol Hill | Weโ€™ve found that while heads-down time at home is fantastic for task-related productivity, in-person time is magic for longer-form productivity. Our in-person days are used to plan, troubleshoot, and check-in with each other on progress and questions. Expect team lunches and whiteboarding.
  • Focused ownership in your role | The rest of the team is here to help you and cares deeply about the long-term functionality of our applications. With that said, weโ€™ll be looking to you to own your lane, go deep, and develop a strong stance on what it takes to make our applications best-in-class.
  • Dedicated monthly AI tooling budget | Weโ€™re in a golden era of AI-powered developer tooling. We strongly encourage augmenting your output with AI tools, and have a dedicated & flexible budget for every team member to support that setup. We care about what you ship, not how.