What Does a Data Engineer Do?

September 3, 2026

(2 min read)

A data engineer builds and operates the systems that move, store, and transform an organisation’s data: the pipelines that collect it from source systems, the warehouse or lakehouse where it lives, and the infrastructure that keeps all of it reliable, governed, and affordable. Every dashboard, report, and machine learning model in the company runs on that foundation. When it works, nobody mentions it; when it breaks, the finance numbers are wrong and nobody can say why.

Most explanations of this role are written for people who want to become data engineers. This one is written for the people who work with them and hire them: what the job actually consists of, where it borders the neighbouring roles, and what separates a senior from a mid-level, because that distinction is where hiring goes wrong.

The work, in detail

The visible output is pipelines: code that extracts data from source systems, transforms it into usable shape, and lands it where analysts, applications, and models can rely on it, on schedule. Around that visible core sits the work that defines the profession.

Modelling and architecture. Deciding how data is structured in the warehouse or lakehouse: Snowflake, Databricks, or BigQuery in most modern stacks, with the transformation layer increasingly in dbt. These decisions are cheap to make and expensive to reverse, which is why they are senior work.

Reliability and observability. Pipelines run unattended at night, against source systems that change without notice. The engineering that matters is defensive: freshness checks on sources, tests on transformations, idempotent tasks that can rerun without double-counting, alerting that fires before the CFO’s dashboard is wrong rather than after.

Streaming, where it is honestly needed. Kafka and its relatives move data in seconds rather than in nightly batches. Real-time is roughly an order of magnitude more operational burden than batch, so a senior data engineer’s contribution is often the question nobody else asks: does this use case actually need it?

Cost discipline. In consumption-priced warehouses, query patterns are budget lines. Part of the job is knowing what the platform costs, which workloads drive it, and what to change when the bill trends wrong.

Governance and lineage. Who may see which data, where a number came from, and proof of both for an auditor. In regulated industries this is not overhead, it is the product.

A realistic week

Career guides describe the role as building pipelines all day. Inside an enterprise team, a senior data engineer’s week is more mixed: reviewing another engineer’s pipeline design before it hardens into production, tracing a data quality incident back to a source system that changed a column meaning without telling anyone, scoping a vague request from an analytics team into something buildable, arguing for (or against) a migration in front of an architecture board, and yes, building. The share of the week spent communicating grows with seniority, which is exactly why hiring processes that only test coding miss what they most need to know.

Where the role sits next to its neighbours

RoleOwnsHandoff with data engineering
Data engineerThe data infrastructure: pipelines, warehouse, reliabilityIs the foundation the others build on
Data scientistModels and analysis built on that dataConsumes clean, reliable data; hands models back for productionising
Software engineerApplications serving usersEmits the source data; pipelines defend against their schema changes
MLOps engineerDeploying and operating ML modelsShares tooling and mindset; owns the model lifecycle rather than the data lifecycle

The data engineer and data scientist border generates the most confusion in hiring, because job ads routinely ask one person to be both. They are different jobs with different instincts, and asking for a single unicorn usually produces a weak version of each.

What separates a senior from a mid-level

Five to six years of production experience is the entry condition; the difference shows in judgement.

A senior owns service levels, not tasks. Ask what freshness the finance pipeline is committed to and what happens when it slips, and a senior answers in specifics, because they have carried that commitment.

A senior designs for the bad day. Backfills, reprocessing, recovery from a corrupted load: mid-level engineers build the happy path and improvise the rest; seniors build the rest first, because they have lived through the improvisation.

A senior negotiates with upstream teams. Data contracts, deprecation notices, staged schema changes. The alternative, silently absorbing whatever source systems do, is how warehouses rot.

A senior can decline to build. The highest-value answer to some requests is a view on an existing table, or the question “what decision will this feed?”. A warehouse with 400 unowned models is what saying yes to everything looks like three years later.

A senior can explain it to a CTO. In enterprise and regulated settings, the engineer who can defend a trade-off in a steering meeting, or walk an auditor through lineage, is worth measurably more than an equally skilled one who cannot. This is why consulting aptitude is one of the four categories in Mahala’s vetting, alongside technical acumen, demonstrated impact, and professional growth.

What “strong on the tools” means at senior level

Tool lists on CVs compress badly. “Kafka” can mean consuming from a topic once, or designing partitioning and consumer groups for a system that cannot lose events. “Snowflake” can mean writing queries, or owning clustering, warehouses sizing, and the monthly bill. “Airflow” can mean writing DAGs, or knowing why a DAG that passes in staging deadlocks on a production backfill. At senior level the tool name means: has operated it in production, has been paged for it, and knows its specific failure modes and cost behaviours. That is the standard a hiring process should test for, and CV keywords cannot carry it.

Which is the honest limit of this article: knowing what the role does tells you what to hire for, not how to verify it. The verification process, CV signals, interview questions, and screening design, is covered in how to hire a data engineer.

Frequently asked questions

Is data engineering the same as ETL development?

ETL, extracting, transforming, and loading data, is one activity inside the role, and historically its centre. The modern role is wider: platform architecture, reliability engineering, cost management, and governance. A team that hires “an ETL developer” for a lakehouse programme is usually under-scoping the job by half.

Do data engineers need to know machine learning?

They need to understand what models consume and how training data gets assembled, because feature pipelines are data engineering work. They do not need to build models; that is the data scientist’s job, and the ML feature work sits on the border the two roles share.

What skills should a job description actually ask for?

SQL and Python as table stakes, one warehouse or lakehouse platform in depth (Snowflake, Databricks, or BigQuery), one orchestrator (Airflow or an equivalent), dbt where the transformation layer uses it, and evidence of production operation: on-call experience, incident stories, cost accountability. Ten-tool wishlists mostly signal that nobody scoped the role.

Where to go from here

What the discipline covers end to end, including how Mahala staffs it, is on our data engineering page. If the reason you looked this up is an empty seat next to a live roadmap, request a vetted shortlist of senior data engineers: two to three blind CVs, production evidence checked, within 72 hours. Prefer to scope the role first? Book a call.

Now you know the role. Meet the people.

Need this role filled rather than explained? Request a vetted shortlist and interview two or three senior Data Engineers in 72 hours.

Featured Articles

Data engineer vs data scientist

2 min read

Data engineer vs data scientist: which do you need?

Data engineer and data scientist sound similar and are constantly confused, including by the recruiters hiring for them. But they are fundamentally different roles, and hiring the wrong one for your problem means it gets solved slowly, expensively, or not at all.
What Is MLOps? A Practical Guide for Enterprises

2 min read

What is MLOps? A practical guide for enterprises

MLOps is the discipline of running machine-learning models in production reliably. If DevOps is what keeps your software alive, MLOps is what keeps your models alive.
Hiring remote specialists in the UAE

2 min read

Hiring remote specialists in the UAE: the efficient default, not the compromise

Hiring senior AI talent in the UAE and across the Gulf is a different problem from hiring in Europe or North America. The ambition is national policy, the procurement is rigorous, and the local pool of specialists.

Request a Vetted Shortlist

About You
About the role
About the engagement

Book a Call