Your responsibilities will include:
- Design, build, and maintain scalable ETL and data ingestion pipelines using Databricks, PySpark, supporting the full lifecycle from development through deployment and optimisation.
- Configure and support Snowplow event tracking and data collection, ensuring high-quality behavioural data is available for analytics and reporting.
- Develop and maintain data transformation and semantic models using SQL and dbt, delivering trusted datasets that enable self-service analytics and business insights.
- Collaborate with software engineers and business stakeholders to understand data requirements, improve data quality, and support integrations across our data platform, including APIs and databases such as PostgreSQL, MongoDB, or DuckDB.
- Contribute to the continuous improvement of our data platform by monitoring pipeline performance, sharing best practices, mentoring colleagues, and exploring new technologies that enhance our data engineering capabilities.
