Seeking a talented and results-oriented Mid-Level Data Engineer to join our growing team. In this pivotal role, you will play a key role in designing, developing, and maintaining our data infrastructure, ensuring the efficient and reliable flow of data to support critical business needs. You will collaborate with data analysts, data scientists, and other stakeholders to translate business requirements into robust data solutions, ultimately driving data-driven decision making across the organization.Job requirements:Role PurposeThe Data Engineer designs, builds, and maintains the data pipelines, platforms, and models that enable reliable, secure, and scalable data flow across the organisation. Reporting to the Head of Data, the incumbent translates the data strategy into hands-on engineering delivery — building modern cloud data platforms, implementing layered (medallion) data architectures, and developing semantic layers that make trusted data easily consumable for reporting, analytics, and AI initiatives.The role combines strong engineering discipline with a practical, business-outcome focus — ensuring that data is accurate, well-modelled, well-governed at the engineering level, and readily accessible to analysts, data scientists, and business stakeholders.Key Responsibilities1. Data Architecture & Cloud Platform EngineeringDesign, build, and maintain scalable cloud data platforms using Google BigQuery and/or Microsoft Fabric.Implement and maintain a medallion (bronze/silver/gold) data architecture, ensuring clear separation between raw, cleansed/conformed, and business-ready data layers.Develop and optimise data warehouse and lakehouse structures, including schemas, partitioning, clustering, and storage strategies for cost and performance.Contribute to enterprise data architecture standards and support migration of legacy data sources into modern cloud platforms.2. Data Pipeline Development & IntegrationDesign, build, and manage robust ELT/ETL pipelines for batch and real-time data processing using industry-standard tools and frameworks.Integrate data from multiple internal and external systems into trusted enterprise data repositories.Automate data ingestion, transformation, and orchestration workflows to reduce manual effort and improve reliability.Apply strong programming skills (Python, SQL) and, where applicable, distributed processing frameworks (e.g. Spark) to build efficient pipelines.3. Semantic Layer & Business Intelligence EnablementDesign, build, and maintain a semantic layer that translates raw and modelled data into consistent, business-friendly metrics, dimensions, and definitions.Ensure a single source of truth for key business metrics, reducing conflicting definitions across reports and tools.Partner with BI, analytics, and data science teams to expose well-modelled datasets through the semantic layer for reporting, dashboarding, and self-service analytics.4. Data Quality, Testing & MonitoringDevelop and implement automated data quality checks, validation rules, and testing across pipelines and data layers.Build proactive monitoring, alerting, and logging to identify and resolve data issues before they impact the business.Troubleshoot data-related incidents, perform root cause analysis, and implement preventative fixes.Maintain version control, CI/CD pipelines, and testing practices for data engineering code.5. Data Security, Access & DocumentationImplement secure data storage and handling practices, applying appropriate access controls at the platform and dataset level.Support compliance with relevant data protection and governance requirements (e.g. POPIA) within engineering processes.Create and maintain clear technical documentation for pipelines, data models, and the semantic layer to support knowledge transfer.Work closely with data analysts, data scientists, and the Head of Data to translate business and analytical requirements into technical specifications.Stay current on emerging data engineering tools and practices, recommending improvements to the organisation's data infrastructure.6. AI Readiness & Advanced Analytics EnablementEnsure data pipelines, data models, and the semantic layer are AI-ready — well-structured, high-quality, and easily consumable by machine learning and AI systems.Prepare and maintain trusted, feature-ready datasets that support machine learning model development, training, and deployment.Enable AI-based data analytics capabilities — such as predictive analytics, natural language querying, and generative AI/BI copilots — by ensuring underlying data is accurate, well-governed, and readily accessible.Collaborate with data scientists and analytics teams to support AI/ML use cases, from data preparation through to production deployment.Stay abreast of emerging AI-driven data engineering practices (e.g. vector databases, retrieval-augmented generation pipelines) and assess their application to the organisation's data platform.