Data engineering platform for synthetic customer engagement pipelines
It provides an end‑to‑end pipeline that ingests synthetic customer data, runs PySpark feature engineering, validates quality, and writes to Delta Lake with idempotent merges. The project includes modular orchestration, transactional outbox delivery, safe replay, and a comprehensive CI suite with 52 tests and high coverage. Designed for data engineers needing a portable, production‑style example that runs on a laptop without private infrastructure. Its thorough testing, observability hooks, and configuration‑driven modules give it robustness beyond typical tutorial code.
View on GitHub →gabrafur/customer-engagement-data-platform