Inkomoko
Senior Data Engineer
Posted
2 weeks ago
Experience
5+ Years
Deadline
Closed
Job Summary
The Senior Data Engineer manages streaming data pipelines, automated testing schedules, storage clusters, dashboard systems, and source code repositories. Core duties include writing data sync pipelines in Python and Airbyte; building streaming tables using Kafka and Debezium; scaling data checks via dbt and Great Expectations; setting up system performance alerts in Grafana; and configuring local virtual machines using Docker.
Key Roles & Architectural Responsibilities
1. Real-Time & Batch Data Pipeline Engineering (30%)
- Architect, scale, and optimize secure, distributed data streaming architectures that power localized business intelligence and real-time operations across five nations.
- Lead the design and deployment of event-driven data streaming pipelines, using Debezium Change Data Capture (CDC) and Apache Kafka to stream transaction logs from core databases and ERP engines.
- Develop, scale, and orchestrate robust batch data ingestion loops using Airbyte and Apache Airflow to pull structural surveys from platforms like SurveyCTO and external web APIs.
- Refine and curate clean, centralized, production-grade datasets to drive advanced analytics, executive reporting dashboards, and cross-border machine learning workloads.
2. Multi-Tier Storage Architecture, Modeling & Data Security (20%)
- Define structural data governance models, writing performant, multi-layered schemas with an emphasis on dimensional modeling (Kimball design paradigms).
- Manage, tune, and maintain high-performance Online Analytical Processing (OLAP) clusters using ClickHouse alongside relational Online Transaction Processing (OLTP) engines running MySQL and PostgreSQL.
- Automate continuous database replication and schema syncing through CDC systems to maintain high availability and instant disaster recovery capabilities.
- Secure regional infrastructure by deploying Role-Based Access Control (RBAC), column-level data masking, end-to-end encryption layers, and automated backup routines.
3. Data Quality Automation & Pipeline Observability (30%)
- Write automated validation tests using dbt and Great Expectations (GX) to catch schema drifts, missing values, or duplicate primary keys before data hits production tables.
- Deploy data anomaly and outlier detection modules, tracking pipeline health and data distribution patterns over time.
- Monitor pipeline performance using Prometheus and Grafana, building custom dashboards to track system metrics—such as processing latency, network throughput, and data freshness—in real time.
4. Infrastructure Automation, DevOps & Mentorship (20%)
- Configure and maintain on-premise bare-metal servers and hypervisor virtual machines (VMs) that host data pipelines, registries, and staging databases.
- Manage CI/CD deployment workflows via Git and Docker containers to ensure identical, repeatable data environment state changes across staging and production.
- Author and update technical reference documents, including data lineage maps, system data dictionaries, and architectural diagrams.
- Coach and mentor junior data engineers, running technical workshops and establishing code review guidelines to upskill the regional engineering team.
Foundational Capabilities & Prerequisites
Required Technical Profile & Experience
- Professional Longevity: A minimum of five (5) or more years of active, hands-on experience in data engineering, backend systems infrastructure, or large-scale analytics database management.
- Academic Credentials: Bachelor's degree in Computer Science, Software Engineering, Data Science, Statistics, or an equivalent engineering discipline (a Master’s degree is considered a distinct asset).
- Scripting & Query Language Mastery: Advanced proficiency in writing clean, object-oriented Python, complex analytical SQL window functions, and native Bash/Shell deployment automation.
- The Modern Data Stack: Proven experience designing and running production systems using tools like dbt, Apache Airflow, Airbyte, ClickHouse, Apache Kafka, and Debezium.
- DevOps Environment Tools: Strong skills configuring virtual machines, version control via Git, and container orchestration utilizing Docker within production settings.
Interpersonal Competencies & Core Values
- Leadership & Trust: A proven ability to mentor junior colleagues, step up to address complex technical debts, and communicate data workflows to non-technical business managers.
- Solutions Orientation: An analytical approach to problem-solving, high attention to detail, and a commitment to data quality standards.
- Mission Alignment: Passion for driving social impact, supporting economic resilience, and working within an inclusive, diverse multicultural environment.
Compensation & Ecosystem Incentives
- Financial Package: Highly competitive regional salary base coupled with potential performance-based annual bonuses.
- Holistic Healthcare Protection: Full medical insurance coverage provided directly for the employee and their immediate dependent family members.
- Retirement Protection: Access to corporate staff savings plans and institutional provident funds, alongside preferred bank rates for long-term employees.
- Recreational Flexibility: Generous annual holiday limits, comprehensive parental leave structures, and dedicated sabbatical options.
- Unparalleled Culture: An inclusive, vibrant, and celebratory team framework that celebrates shared wins and supports teams through challenges (We Eat Goat!).
Skills Required:
- Computer / Software / It / Data
- Economics / Statistics
Quick Actions
Share Vacancy