Mastronardi Produce pioneered the commercial greenhouse industry in North America, and we’re now the leading greenhouse vegetable company on the continent. Our award-winning, flavorful produce is packed under the SUNSET® brand and is available at leading grocery retailers across North America. Family owned for over 70 years, we pride ourselves on having the most flavorful products and the best people in the industry. We are constantly pushing boundaries to be a leader in fresh produce innovation. We seek individuals that demonstrate our PRIDE values (Passion, Respect, Innovation, Drive, Excellence) to help us fulfill our mission to inspire healthy living through WOW flavor experiences.
Our head office in Kingsville, ON is currently seeking a Data Warehouse Engineer to join our team! As a Data Warehouse Engineer, you will help enhance and maintain Mastronardi Produce’s One Data Platform (ODP), our enterprise data foundation built on Microsoft Fabric using Medallion (Bronze/Silver/Gold) architecture. Most development is done in PySpark notebooks, so strong coding skills are essential. You will profile and analyze data to meet stakeholder requirements, build and optimize pipelines, and design data models that are both technically sound and easy for end users to consume or to build their own reports. This role requires strong SQL, Spark/Python, and data modeling craft, sound engineering judgment across the full pipeline lifecycle, and the ability to guide and review the work of other data engineers.
Values:
To perform the job successfully, the incumbent’s behavior must be consistent with the PRIDE values expected of all Mastronardi Produce employees: be Passionate; have Respect; be Innovative; be Driven and strive for Excellence.
Primary Responsibilities:
- Data modeling & stakeholder requirements — Profile and analyze data to translate business and reporting requirements into performant, well-modeled Gold-layer datasets.
- New data source onboarding — Profile new data sources and determine the appropriate ingestion method and refresh approach (full load, incremental, CDC, or near-real-time) based on source system constraints, data volume, and business need.
- Pipeline development & optimization — Design, build, and continuously optimize data pipelines, with a focus on reducing pipeline run times and improving reliability at scale.
- Troubleshooting & root cause resolution — Own the investigation, root-causing, and resolution of data engineering-related pipeline and data issues, partnering with source system owners where needed.
- Self-service data foundation — Build and evolve a data foundation that is flexible and easy for end users to work with directly — enabling them to write their own SQL and build their own reports against governed, well-documented Gold-layer models.
- Partnership with report developers — Work closely with report developers throughout the build process to ensure requirements are met and that final data is accurate, well-modeled, and performant.
- Partnership with business stakeholders — Partner with stakeholders across Finance, Supply Chain, Sales, Operations, and Logistics to understand data requirements and business logic, and translate them into scalable data models.
- Guidance to other engineers — Review the work of other data engineers and provide guidance on approach, technique, and best practices, helping raise the overall quality and consistency of engineering across ODP.
- Holistic platform thinking — Think across projects and domains — rather than in isolation — to build and enhance a data foundation that scales as one cohesive platform instead of disconnected, duplicative solutions.
- Governance & documentation — Implement data governance, security, and access control practices consistent with ODP standards, and maintain clear documentation for data models, pipelines, and workflows.
- Team orientation — Operate as a team player who executes on what’s best for the business, collaborating effectively with engineers, report developers, and stakeholders in a fast-paced, evolving environment.
Education/Background Requirements:
- Bachelor’s degree in computer science, information systems, or a related field.
- MS Fabric Platform: Strong Hands-on experience with Microsoft Fabric (Lakehouses, Data Pipelines/Dataflows, Notebooks) and strong working knowledge of Medallion architecture (Bronze/Silver/Gold layer design).
- Other Cloud Data Platforms: Good hands-on experience with Synapse, Amazon, S3, Azure Data Lake Storage or Google Cloud Storage.
- SQL Expertise: Expert-level SQL, including complex joins, window functions, and query performance tuning, with the ability to design SQL-first data models that business users and report developers can query directly and confidently
- Data Modeling: Strong data modeling skills, including dimensional (star schema) modeling, grain definition, and slowly changing dimensions, with an emphasis on designing schemas that are both engineering-sound and easy for end users to navigate.
- Spark & Python: Expert-level PySpark and Python, as most Fabric development is done in Spark notebooks — including building and optimizing transformations, and writing clean, maintainable, production-grade notebook code.
- Data Source Profiling: Demonstrated ability to profile unfamiliar data sources and recommend the right ingestion and refresh strategy based on source constraints, volume, and business need.
- Pipeline Optimization: Experience identifying performance bottlenecks and reducing pipeline run times in a production data engineering environment.
- Troubleshooting: Strong root-cause analysis and troubleshooting skills for data quality and pipeline issues.
- Stakeholder & Cross-functional Partnership: Experience gathering requirements directly from business stakeholders and partnering with report developers to validate data accuracy and performance.
- Peer Guidance: Experience reviewing others’ technical work and providing constructive, actionable guidance, even without formal management authority.
Specific Knowledge, Skills and Abilities Required
- Broader platform exposure: Experience with other big data or cloud platforms (e.g., Hadoop, Databricks, AWS, Azure, or GCP) is a plus, though the core platform for this role is Microsoft Fabric.
- AI-assisted Development: Experience using AI coding assistants and LLM-based tools (e.g., Claude, GitHub Copilot) to accelerate data engineering work, writing and reviewing SQL/PySpark, debugging pipelines, and generating documentation while retaining full ownership of correctness, performance, and quality.
- Real-time Processing: Familiarity with real-time/streaming data processing (e.g., Kafka, Flink, Fabric Eventstream) is a plus. Note: near-real-time pipelines have not yet been built on ODP due to source system constraints and competing priorities, so this is a forward-looking skill rather than a day-one requirement.
- Governance Tooling: Familiarity with data governance frameworks, data quality management, and metadata/catalog tools.
- Tooling & Process: Experience with version control (GIT) and agile development practices.
- Certifications in relevant technologies (e.g., Fabric Analytics Engineer. AWS Certified Big Data Specialty) are a plus.
Working Conditions:
- Typical office environment.
Please note: Mastronardi Produce has accommodation processes and policies in place and provides accommodation for employees with disabilities. If you require a specific accommodation because of a disability or documented medical need, please contact the Human Resource office so that arrangements can be made for the appropriate accommodation to be put into place.
Salary is $85k/yr-$95k/yr CAD
“Mastronardi uses AI to source and select candidates on our job platforms.”