Senior Data Engineer / Data Scientist – Azure, Databricks, Statistical Modelling
Location: Toronto, Ontario
Work Model: Hybrid – 2 days onsite per week
Employment Type: Contract
Duration: 6 months, with potential extension
Pay Rate: CAD 71–77/hour
Experience Level: Senior
Public-Sector Experience: Required
Role Overview
We are seeking a senior-level Data Engineer / Data Scientist who can support both modern cloud data-platform development and advanced statistical analytics.
The successful candidate will design and build Azure-based data pipelines, data lakes and Lakehouse solutions while also analyzing, validating and operationalizing statistical, predictive and analytical models.
This role requires someone who can work across data engineering, data science, analytics, business requirements and stakeholder consulting. The candidate must be comfortable working directly with data engineers, data architects, analysts, researchers, technical teams, data stewards and business stakeholders.
Key Responsibilities
- Design, develop and maintain cloud-based data lakes, Lakehouse environments and automated data pipelines.
- Build ingestion, transformation, validation and data-quality workflows for structured and unstructured data.
- Develop scalable ETL and ELT pipelines using Azure Data Factory, Azure Databricks, Azure Data Lake and Azure Synapse.
- Implement Lakehouse and medallion architecture using Bronze, Silver and Gold data layers.
- Develop and maintain analytical datasets, data models, dashboards and reporting solutions.
- Analyze existing statistical, forecasting, simulation or research models and determine their data, technology and implementation requirements.
- Build, validate, troubleshoot and operationalize statistical and analytical models using R and/or Python.
- Assess model logic, assumptions, parameters, data dependencies, outputs and limitations.
- Perform model validation, sensitivity analysis, calibration and data-quality assessment.
- Develop reproducible analytical workflows using version-controlled code.
- Support machine-learning, natural-language-processing, forecasting and scenario-analysis use cases.
- Adapt externally developed, academic or research-based analytical models for use within an enterprise production environment.
- Translate business, research and analytical requirements into sustainable technical solutions.
- Assess client needs, objectives, constraints and available implementation options.
- Troubleshoot complex failures involving data pipelines, cloud services, analytical code, models and downstream reporting.
- Create visualizations, dashboards and reports using Power BI or comparable tools.
- Support data acquisition, data governance, privacy, security and data-sharing activities.
- Prepare technical documentation, model documentation, data dictionaries, implementation guides and operational procedures.
- Lead knowledge-transfer sessions and train technical and non-technical staff.
- Work collaboratively with data architects, ETL developers, data scientists, analysts, business teams and client stakeholders.
Mandatory Qualifications
- Senior-level experience in data engineering, data science, statistical analytics or a closely related field.
- Hands-on experience building and supporting end-to-end cloud-based data and analytics solutions.
- Strong experience with Microsoft Azure data technologies, including:
- Azure Data Lake Storage
- Azure Databricks
- Azure Data Factory
- Azure Synapse Analytics
- Azure Storage
- Strong programming experience using Python and/or R.
- Strong SQL skills, including complex queries, data transformations and analytical data preparation.
- Experience designing and implementing cloud data lakes, Lakehouse solutions and medallion architecture.
- Experience developing automated ingestion, ETL/ELT, transformation and data-quality pipelines.
- Experience developing reproducible and version-controlled analytical workflows using Git or similar tools.
- Senior-level knowledge of statistical methods, analytical models and their underlying data requirements.
- Experience validating models, testing assumptions and resolving data or code issues.
- Experience developing or supporting machine-learning, forecasting, simulation, NLP or predictive-modelling solutions.
- Experience translating business or research requirements into technical and analytical solutions.
- Strong troubleshooting and root-cause-analysis skills across complex, multi-component environments.
- Experience preparing technical documentation and conducting knowledge-transfer sessions.
- Strong consulting, stakeholder-management and client-facing communication skills.
- Ability to explain complex analytical methods, assumptions, findings and limitations to technical and non-technical audiences.
- Previous Canadian public-sector experience.
- Ability to work onsite in Toronto two days per week.
Preferred Qualifications
- Experience with Power BI and enterprise dashboard development.
- Experience with PySpark, Apache Spark and Delta Lake.
- Experience with Azure Machine Learning, MLflow or MLOps.
- Experience with Unity Catalog, Databricks Workflows or Databricks Auto Loader.
- Experience with data governance, lineage, metadata management and quality controls.
- Experience working with healthcare, public-health, population-health or government data.
- Experience with chronic-disease, burden-of-disease, epidemiological, microsimulation or population-modelling initiatives.
- Experience operationalizing an academic, external or research-based model within an enterprise environment.
- Experience with model calibration, sensitivity analysis, uncertainty analysis or scenario modelling.
- Experience with RStudio, R Markdown, Quarto, Shiny or reproducible research practices.
- Familiarity with accessibility standards, including AODA.
- Experience leading or mentoring data engineers, analysts or data scientists.
Technical Environment
- Microsoft Azure
- Azure Data Lake Storage
- Azure Databricks
- Azure Data Factory
- Azure Synapse Analytics
- Python
- R
- SQL
- PySpark
- Spark
- Delta Lake
- Power BI
- Git
- Azure DevOps
- Machine learning
- Statistical modelling
- Forecasting
- NLP
- Data governance
- Lakehouse and medallion architecture
Pay: $70,440.45-$152,634.69 per year
Work Location: Hybrid remote in Toronto, ON (York District)