The Baldwin Group · 2022–2026

Data Lake buildout on AWS using Databricks

Dev, Stage and Prod Databricks workspaces on AWS with Unity Catalog RBAC, federated sources, Lakebase apps and shared AI coding standards for the team.

AWS · Databricks · MLflow · Terraform

Problem

Source data sat in operational systems, models lived in notebooks, lead scoring ran on a paid vendor model, and deployments were manual.

Work

  • Dev, Stage and Prod Databricks workspaces on AWS, provisioned in Terraform: networking, storage and cross-account IAM.
  • RBAC over all data through Unity Catalog groups, schemas and grants.
  • External connections to Snowflake, Microsoft Fabric, DynamoDB and Postgres.
  • 27 Delta Live Tables pipelines, deployed with GitHub Actions and Databricks Asset Bundles; 180 TB of unstructured data cataloged in Unity Catalog.
  • Batch and online inference on MLflow Model Registry, Feature Store and Delta Lake, validated on held-out data before promotion.
  • Applications built on Lakebase.
  • Agentic engineering with Genie Code: shipped team skills to every engineer’s AI coding harness, including contractors and consultants across time zones, to enforce pull request standards.
  • Pandas-to-PySpark rewrite of a finance pipeline: 55 min to 15 min.
  • A pipeline template (incremental and batch loads, SCD Type I and II) covering about 80% of new pipelines.

Result

  • One governed platform across three environments, with access controlled in Unity Catalog.
  • An in-house model replaced the vendor model and doubled lead-scoring accuracy (0.4 to 0.8).
  • The same review standards applied to every engineer and every AI assistant, regardless of time zone.

← All work