A data engineer builds and operates the systems that move, store, and transform an organisation’s data: the pipelines that collect it from source systems, the warehouse or lakehouse where it lives, and the infrastructure that keeps all of it reliable, governed, and affordable. Every dashboard, report, and machine learning model in the company runs on that foundation. When it works, nobody mentions it; when it breaks, the finance numbers are wrong and nobody can say why.
Most explanations of this role are written for people who want to become data engineers. This one is written for the people who work with them and hire them: what the job actually consists of, where it borders the neighbouring roles, and what separates a senior from a mid-level, because that distinction is where hiring goes wrong.
The work, in detail
The visible output is pipelines: code that extracts data from source systems, transforms it into usable shape, and lands it where analysts, applications, and models can rely on it, on schedule. Around that visible core sits the work that defines the profession.
Modelling and architecture. Deciding how data is structured in the warehouse or lakehouse: Snowflake, Databricks, or BigQuery in most modern stacks, with the transformation layer increasingly in dbt. These decisions are cheap to make and expensive to reverse, which is why they are senior work.
Reliability and observability. Pipelines run unattended at night, against source systems that change without notice. The engineering that matters is defensive: freshness checks on sources, tests on transformations, idempotent tasks that can rerun without double-counting, alerting that fires before the CFO’s dashboard is wrong rather than after.
Streaming, where it is honestly needed. Kafka and its relatives move data in seconds rather than in nightly batches. Real-time is roughly an order of magnitude more operational burden than batch, so a senior data engineer’s contribution is often the question nobody else asks: does this use case actually need it?
Cost discipline. In consumption-priced warehouses, query patterns are budget lines. Part of the job is knowing what the platform costs, which workloads drive it, and what to change when the bill trends wrong.
Governance and lineage. Who may see which data, where a number came from, and proof of both for an auditor. In regulated industries this is not overhead, it is the product.
A realistic week
Career guides describe the role as building pipelines all day. Inside an enterprise team, a senior data engineer’s week is more mixed: reviewing another engineer’s pipeline design before it hardens into production, tracing a data quality incident back to a source system that changed a column meaning without telling anyone, scoping a vague request from an analytics team into something buildable, arguing for (or against) a migration in front of an architecture board, and yes, building. The share of the week spent communicating grows with seniority, which is exactly why hiring processes that only test coding miss what they most need to know.
Where the role sits next to its neighbours
| Role | Owns | Handoff with data engineering |
|---|---|---|
| Data engineer | The data infrastructure: pipelines, warehouse, reliability | Is the foundation the others build on |
| Data scientist | Models and analysis built on that data | Consumes clean, reliable data; hands models back for productionising |
| Software engineer | Applications serving users | Emits the source data; pipelines defend against their schema changes |
| MLOps engineer | Deploying and operating ML models | Shares tooling and mindset; owns the model lifecycle rather than the data lifecycle |
The data engineer and data scientist border generates the most confusion in hiring, because job ads routinely ask one person to be both. They are different jobs with different instincts, and asking for a single unicorn usually produces a weak version of each.
What separates a senior from a mid-level
Five to six years of production experience is the entry condition; the difference shows in judgement.
A senior owns service levels, not tasks. Ask what freshness the finance pipeline is committed to and what happens when it slips, and a senior answers in specifics, because they have carried that commitment.
A senior designs for the bad day. Backfills, reprocessing, recovery from a corrupted load: mid-level engineers build the happy path and improvise the rest; seniors build the rest first, because they have lived through the improvisation.
A senior negotiates with upstream teams. Data contracts, deprecation notices, staged schema changes. The alternative, silently absorbing whatever source systems do, is how warehouses rot.
A senior can decline to build. The highest-value answer to some requests is a view on an existing table, or the question “what decision will this feed?”. A warehouse with 400 unowned models is what saying yes to everything looks like three years later.
A senior can explain it to a CTO. In enterprise and regulated settings, the engineer who can defend a trade-off in a steering meeting, or walk an auditor through lineage, is worth measurably more than an equally skilled one who cannot. This is why consulting aptitude is one of the four categories in Mahala’s vetting, alongside technical acumen, demonstrated impact, and professional growth.
What “strong on the tools” means at senior level
Tool lists on CVs compress badly. “Kafka” can mean consuming from a topic once, or designing partitioning and consumer groups for a system that cannot lose events. “Snowflake” can mean writing queries, or owning clustering, warehouses sizing, and the monthly bill. “Airflow” can mean writing DAGs, or knowing why a DAG that passes in staging deadlocks on a production backfill. At senior level the tool name means: has operated it in production, has been paged for it, and knows its specific failure modes and cost behaviours. That is the standard a hiring process should test for, and CV keywords cannot carry it.
Which is the honest limit of this article: knowing what the role does tells you what to hire for, not how to verify it. The verification process, CV signals, interview questions, and screening design, is covered in how to hire a data engineer.
Frequently asked questions
Is data engineering the same as ETL development?
ETL, extracting, transforming, and loading data, is one activity inside the role, and historically its centre. The modern role is wider: platform architecture, reliability engineering, cost management, and governance. A team that hires “an ETL developer” for a lakehouse programme is usually under-scoping the job by half.
Do data engineers need to know machine learning?
They need to understand what models consume and how training data gets assembled, because feature pipelines are data engineering work. They do not need to build models; that is the data scientist’s job, and the ML feature work sits on the border the two roles share.
What skills should a job description actually ask for?
SQL and Python as table stakes, one warehouse or lakehouse platform in depth (Snowflake, Databricks, or BigQuery), one orchestrator (Airflow or an equivalent), dbt where the transformation layer uses it, and evidence of production operation: on-call experience, incident stories, cost accountability. Ten-tool wishlists mostly signal that nobody scoped the role.
Where to go from here
What the discipline covers end to end, including how Mahala staffs it, is on our data engineering page. If the reason you looked this up is an empty seat next to a live roadmap, request a vetted shortlist of senior data engineers: two to three blind CVs, production evidence checked, within 72 hours. Prefer to scope the role first? Book a call.