About the position
Qualifications:
- Bachelor’s degree – Computer Science, Data Science, Information Technology, or related field.
- Relevant certifications (e.g. TOGAF, CDMP, DAMA-DMBOK).
- Microsoft Azure experience, training and certification will be advantageous.
- Databricks experience, training and certification will be advantageous.
Deliverables:
- Python orchestration.
- SQL data modelling.
- Databricks SDK automation.
- CI/CD via Azure DevOps.
- Gather, organize and analyze data from various databases, source systems or external systems to profile, monitor and evaluate data quality.
- Implement data quality workflows, mappings, mapplets, analyst profiles, scorecards and reference tables and automation thereof. Develops data quality key performance indicators (KPI’s) and reporting to measure, monitor and evaluate data quality across the entire data stack.
- Develop a DQ dashboard and mappings for business to track the quality of their data.
- Track, monitor and document testing results post DQ rule set implementation.
- Performs root cause analysis and collaborates with data stakeholders to identify and understand factors that contribute to data quality issues.
- Recommends data capture and operational process improvements based on findings from RCA’s.
- Coordinates with the appropriate internal (e.g. IT, Operations) and external stakeholders (e.g. technology or data partners) in the correction of source data or creation of translation sources, where applicable.
- The development and maintenance of Extract Transform and Load (ETL) processes, database and performance administration, and dimensional design of the table structure. The Data Integration Engineer will need to engage with several stakeholders including Project Managers, Multiple Vendor resources for delivery, the company’s Commercial Business Units and EIS Management.
Technology Stacks:
- Platform: Azure Databricks (Serverless + SQL Warehouses + Instance Pools)
- Languages: Python (primary), SQL (heavy — aggregation scripts), some Trino SQL
- Data: Unity Catalog, Delta Lake
- SDKs: Databricks SDK, databricks-SQL-connector, Pandas
- Orchestration: Lakeflow Jobs (parameterised multi-task DAGs, task values, run_job_task chaining)
- Source Control: Azure DevOps Git + Azure Pipelines CI/CD
- Cross-platform: Trino connectivity for hybrid queries
- Patterns: Parallelism, parameterisation, retention management, medallion-like layering
Qualifications:
- Bachelor’s degree – Computer Science, Data Science, Information Technology, or related field.
- Relevant certifications (e.g. TOGAF, CDMP, DAMA-DMBOK).
- Microsoft Azure experience, training and certification will be advantageous.
- Databricks experience, training and certification will be advantageous.
Deliverables:
- Python orchestration.
- SQL data modelling.
- Databricks SDK automation.
- CI/CD via Azure DevOps.
- Gather, organize and analyze data from various databases, source systems or external systems to profile, monitor and evaluate data quality.
- Implement data quality workflows, mappings, mapplets, analyst profiles, scorecards and reference tables and automation thereof. Develops data quality key performance indicators (KPI’s) and reporting to measure, monitor and evaluate data quality across the entire data stack.
- Develop a DQ dashboard and mappings for business to track the quality of their data.
- Track, monitor and document testing results post DQ rule set implementation.
- Performs root cause analysis and collaborates with data stakeholders to identify and understand factors that contribute to data quality issues.
- Recommends data capture and operational process improvements based on findings from RCA’s.
- Coordinates with the appropriate internal (e.g. IT, Operations) and external stakeholders (e.g. technology or data partners) in the correction of source data or creation of translation sources, where applicable.
- The development and maintenance of Extract Transform and Load (ETL) processes, database and performance administration, and dimensional design of the table structure. The Data Integration Engineer will need to engage with several stakeholders including Project Managers, Multiple Vendor resources for delivery, the company’s Commercial Business Units and EIS Management.
Technology Stacks:
- Platform: Azure Databricks (Serverless + SQL Warehouses + Instance Pools)
- Languages: Python (primary), SQL (heavy — aggregation scripts), some Trino SQL
- Data: Unity Catalog, Delta Lake
- SDKs: Databricks SDK, databricks-SQL-connector, Pandas
- Orchestration: Lakeflow Jobs (parameterised multi-task DAGs, task values, run_job_task chaining)
- Source Control: Azure DevOps Git + Azure Pipelines CI/CD
- Cross-platform: Trino connectivity for hybrid queries
- Patterns: Parallelism, parameterisation, retention management, medallion-like layering
Desired Skills:
- Degree
- TOGAF
- CDMP
- DAMA-DMBOK
- Python orchestration.
- SQL data modelling.
- Databricks SDK automation.
- CI/CD via Azure DevOps.