31 - Data Engineer
Build and operate the data pipelines behind a large European enterprise's analytics platform. Python, PySpark, Airflow and AWS, on live production systems.
Our client is looking for a Data Engineer to join a 10-person team building and operating the data platform behind a large European enterprise's analytical applications.
This is a five-month engagement.
The work sits on production systems. The platform is already live and feeding analytical applications that people rely on, and it is still evolving — new sources land, pipelines get reshaped, infrastructure gets rebuilt underneath you. You'll be designing and operating pipelines in that environment rather than building something greenfield in isolation.
This is an individual-contributor role. You won't be managing anyone, and you won't be handed fully specified tickets either. The team expects you to take a data problem, work out how the pipeline should be shaped, build it, and then own it in production.
Mid-level here means you've already done this work somewhere real. You've had a pipeline break at an inconvenient hour and you've fixed it. You know what a backfill costs. You don't need someone standing over you to get a PySpark job into production.
What you'll be doing
Designing and building data pipelines that feed the client's analytical applications
Operating and maintaining existing production pipelines — monitoring, debugging, fixing, improving
Writing and optimising distributed processing jobs in PySpark
Building and maintaining orchestration in Airflow: DAG design, dependencies, scheduling, retries, backfills
Developing ETL workflows from source systems into the analytical layer
Modelling and writing SQL against PostgreSQL for analytical workloads
Contributing to the continuous evolution of the client's cloud-based (AWS) data infrastructure
Working within existing CI/CD pipelines and Git workflows — reviewed code, tested changes, no direct-to-production edits
Investigating data quality issues and making pipelines more resilient to the ones that keep coming back
Collaborating with the wider platform and analytics teams on what the data needs to look like downstream
Required skills and experience
Python — strong, production-level. This is the primary language of the role.
PySpark — hands-on experience writing and tuning distributed processing jobs
Airflow — you've built and operated DAGs, not just triggered someone else's
AWS — practical experience with cloud-based data infrastructure
PostgreSQL and SQL — comfortable writing and reasoning about analytical queries
ETL design — you can take a source system and a target and work out the pipeline in between
CI/CD and Git — you work inside a pipeline, with branches, reviews and automated checks
Experience on live production data systems — not only development environments, academic projects or course work
Comfortable in a fast-paced delivery environment where the platform changes while you're working on it
Professional English, spoken and written, and the communication habits that make fully remote work function
Nice to have
Hadoop ecosystem experience
Data modelling for analytics and BI consumption
Data quality, monitoring or observability tooling
Exposure to infrastructure as code or containerised workloads
Experience working with an enterprise client or in a regulated environment
- Department
- Outstaffing
- Remote status
- Fully Remote
About Tunga
Tunga is the go-to platform for hiring African software developers. Companies from all over the world use Tunga to hire African software developers to execute software projects, as full-time or part-time members of distributed software teams.
Tunga’s mission is to create tech jobs for African youths and has a community of over 3000 software developers.
We were founded in 2015 and have served over 250 clients from all over the world. Tunga’s clients have a diverse profile: SMEs, startups, corporates, and NGOs all belong to our client base.