About the position
Fully Remote Working
Role purpose and context:
Data science is not a support function at the company — it is central to the product. The company privacy-preserving models, entity-resolution algorithms, data-linkage methods, analytics, and AI-augmented pipelines are core to the value the company delivers to clients.
The Lead Data Scientist serves as second-in-command to the Head of Data Science and combines senior technical leadership with hands-on delivery. The role requires someone who can lead complex analytical work, represent the Data Science function when required, and remain actively involved in designing, coding, validating, and delivering solutions.
You will also play a key role in company transition from analytics built primarily through traditional SQL, Synapse Pipelines, and Power BI toward Python-powered Jupyter workflows that use Claude Code and other LLM tools to accelerate research, coding, documentation, prototyping, and production delivery.
The environment includes PySpark, Delta Lake, Jupyter Notebooks, Azure Synapse, Trino, Power BI, Kubernetes-hosted model serving, and Fast API inference endpoints.
Engagement and location:
- Location: South Africa (Remote) Permanent
- Occasional travel (a few times/year – in person meeting/s) to Johannesburg or Cape Town required; proximity to either city preferred.
- Reports to: Head of Data Science.
- May lead an Agile delivery squad within the Platform organization.
Key responsibilities:
- Own the analytical roadmap. Help shape and prioritise the roadmap and backlog so analytical work is clearly defined, developed to a high standard, and delivered on time. Maintain a coherent customer analytics journey while identifying opportunities to create reusable assets.
- Own customer analytical delivery end to end. Translate ambiguous business questions into decision-ready analysis by framing the problem, defining success criteria, conducting the analysis, and presenting findings to senior stakeholders. Be prepared to explain clearly when the evidence does not support the customer's expected conclusion.
- Write production-grade Python. Develop maintainable, testable code using modern software-engineering practices and toolchains.
- Apply warehouse and Lakehouse best practices. Reason confidently about data grain, joins, provenance, effective dating, reconciliation, and traceability so published results can be linked back to source data and analytical conclusions reflect the quality and precision of that data.
- Use AI throughout the delivery lifecycle. Apply Claude Code and agentic workflows to research, prototyping, development, testing, documentation, and production delivery.
- Generate insight across multiple data parties. Work with matched, indirect, aggregated, or otherwise imperfect datasets using techniques such as weighted binning, dependency and importance measures, and validation controls. Understand the limitations of the data and recognise invalid or misleading comparisons.
- Support Data Science leadership. Act as deputy to the Head of Data Science when required, mentor team members, support delivery management, and represent the function with internal and external stakeholders.
Must-have technical skills / experience:
- Senior-level data science expertise across machine learning, statistical modelling, and AI system design in production environments, including deploying, monitoring, and retraining models.
- Advanced Python skills and strong experience with Jupyter Notebooks, pandas, NumPy, and scikit-learn.
- Experience with Delta Lake or comparable Lakehouse architectures, including PySpark, feature stores, training-data versioning, and schema evolution.
- Strong data warehouse expertise, including Azure Synapse Analytics, Azure SQL, T-SQL, and/or ANSI SQL.
- Strong analytics and visualisation experience using Power BI and Python-based visualisation or interactive-analysis tools.
- Team-level technical or engineering leadership experience, including mentoring data scientists, owning delivery, and communicating with senior stakeholders.
- Excellent written and verbal English communication skills.
Minimum experience:
- 7+ years of experience across data science, data analytics, or data warehousing, with substantial recent data science experience.
- 2+ years in a technical or engineering leadership role.
- Bachelor's degree in a STEM discipline or equivalent professional experience.
Desirable experience:
- Apache Ranger or Comparable data-governance technologies
- Fast API
- Automated testing with PyTest
- PySpark and distributed data processing.
- Kubernetes-based model deployment or serving.
- C#.NET development
- Angular or Electron development
Why you’ll love working for the company:
The company believe in taking care of their team and creating an environment where you can thrive. As part of their company, you’ll enjoy:
- Flexible working arrangements: Whether you’re a night owl or an early bird, they offer hybrid and remote options to suit your lifestyle
- Comprehensive benefits: From a wellness program to home office reimbursements and continuous learning opportunities, they have got you covered.
- Team culture: Fun team-building activities, regular socials, and a supportive, inclusive culture that values transparency, accountability, and work-life balance.
Performance incentives: Competitive salaries, ESOP, and recognition for your hard work
Fully Remote Working
Role purpose and context:
Data science is not a support function at the company — it is central to the product. The company privacy-preserving models, entity-resolution algorithms, data-linkage methods, analytics, and AI-augmented pipelines are core to the value the company delivers to clients.
The Lead Data Scientist serves as second-in-command to the Head of Data Science and combines senior technical leadership with hands-on delivery. The role requires someone who can lead complex analytical work, represent the Data Science function when required, and remain actively involved in designing, coding, validating, and delivering solutions.
You will also play a key role in company transition from analytics built primarily through traditional SQL, Synapse Pipelines, and Power BI toward Python-powered Jupyter workflows that use Claude Code and other LLM tools to accelerate research, coding, documentation, prototyping, and production delivery.
The environment includes PySpark, Delta Lake, Jupyter Notebooks, Azure Synapse, Trino, Power BI, Kubernetes-hosted model serving, and Fast API inference endpoints.
Engagement and location:
- Location: South Africa (Remote) Permanent
- Occasional travel (a few times/year – in person meeting/s) to Johannesburg or Cape Town required; proximity to either city preferred.
- Reports to: Head of Data Science.
- May lead an Agile delivery squad within the Platform organization.
Key responsibilities:
- Own the analytical roadmap. Help shape and prioritise the roadmap and backlog so analytical work is clearly defined, developed to a high standard, and delivered on time. Maintain a coherent customer analytics journey while identifying opportunities to create reusable assets.
- Own customer analytical delivery end to end. Translate ambiguous business questions into decision-ready analysis by framing the problem, defining success criteria, conducting the analysis, and presenting findings to senior stakeholders. Be prepared to explain clearly when the evidence does not support the customer's expected conclusion.
- Write production-grade Python. Develop maintainable, testable code using modern software-engineering practices and toolchains.
- Apply warehouse and Lakehouse best practices. Reason confidently about data grain, joins, provenance, effective dating, reconciliation, and traceability so published results can be linked back to source data and analytical conclusions reflect the quality and precision of that data.
- Use AI throughout the delivery lifecycle. Apply Claude Code and agentic workflows to research, prototyping, development, testing, documentation, and production delivery.
- Generate insight across multiple data parties. Work with matched, indirect, aggregated, or otherwise imperfect datasets using techniques such as weighted binning, dependency and importance measures, and validation controls. Understand the limitations of the data and recognise invalid or misleading comparisons.
- Support Data Science leadership. Act as deputy to the Head of Data Science when required, mentor team members, support delivery management, and represent the function with internal and external stakeholders.
Must-have technical skills / experience:
- Senior-level data science expertise across machine learning, statistical modelling, and AI system design in production environments, including deploying, monitoring, and retraining models.
- Advanced Python skills and strong experience with Jupyter Notebooks, pandas, NumPy, and scikit-learn.
- Experience with Delta Lake or comparable Lakehouse architectures, including PySpark, feature stores, training-data versioning, and schema evolution.
- Strong data warehouse expertise, including Azure Synapse Analytics, Azure SQL, T-SQL, and/or ANSI SQL.
- Strong analytics and visualisation experience using Power BI and Python-based visualisation or interactive-analysis tools.
- Team-level technical or engineering leadership experience, including mentoring data scientists, owning delivery, and communicating with senior stakeholders.
- Excellent written and verbal English communication skills.
Minimum experience:
- 7+ years of experience across data science, data analytics, or data warehousing, with substantial recent data science experience.
- 2+ years in a technical or engineering leadership role.
- Bachelor's degree in a STEM discipline or equivalent professional experience.
Desirable experience:
- Apache Ranger or Comparable data-governance technologies
- Fast API
- Automated testing with PyTest
- PySpark and distributed data processing.
- Kubernetes-based model deployment or serving.
- C#.NET development
- Angular or Electron development
Why you’ll love working for the company:
The company believe in taking care of their team and creating an environment where you can thrive. As part of their company, you’ll enjoy:
- Flexible working arrangements: Whether you’re a night owl or an early bird, they offer hybrid and remote options to suit your lifestyle
- Comprehensive benefits: From a wellness program to home office reimbursements and continuous learning opportunities, they have got you covered.
- Team culture: Fun team-building activities, regular socials, and a supportive, incl