Data you can trace back to the source
Our data engineering services move data out of spreadsheets and legacy databases into tested, documented pipelines. Every table has an owner, a definition and written access rules, so your reports and models agree.
- 1–3 week data audit to start
- Quality tests on every load
- Owners and access rules documented
Pipelines, storage and the rules that keep data trustworthy
Good data engineering is mostly unglamorous: knowing where each number comes from, catching bad loads before anyone reports on them, and not paying for compute nobody uses. That's the work we do.
Batch and streaming pipelines
Extraction from application databases, SaaS tools, files and APIs, with incremental loads, retries and backfills. Streaming with Pub/Sub, Kinesis or Kafka where a decision really needs to happen in seconds.
Warehouse and lakehouse design
Postgres or Supabase for smaller workloads, BigQuery or Redshift as a cloud warehouse, or object storage with an open table format when you need to keep raw history cheaply.
Data modelling
Raw, cleaned and business layers, with dimensional models that match how your teams ask questions. Definitions such as "active customer" live in one place, in version control.
Quality checks and data contracts
Tests for nulls, duplicates, volumes and freshness on every load, plus written contracts with the teams that own source systems, so a renamed column breaks a test rather than a board report.
Orchestration
Scheduled, dependency-aware jobs with alerting and clear ownership, using Airflow, Dagster or your cloud's native scheduler, depending on what your team will maintain.
Governance and access control
Role-based access, column-level restrictions for personal data, audit logs and a data catalogue that says who owns each table and what it means.
Do you need a warehouse, a lakehouse, or just Postgres?
Start with Postgres if your analytical data is modest and a handful of people query it. Move to a cloud warehouse when queries slow down or analysts multiply. Add a lakehouse layer when you must keep large volumes of raw or semi-structured data, such as logs, events or documents, for ML at low storage cost.
| Option | Fits | Typical stack | Watch out for |
|---|---|---|---|
| Postgres or Supabase | Early-stage products, internal reporting, data that's mostly structured and measured in gigabytes | Postgres with a separate reporting schema or read replica, pgvector for embeddings | Heavy analytical queries competing with the application. Plan the move before it hurts. |
| Cloud warehouse | Many analysts, BI tools, larger history, joins across many sources | BigQuery on Google Cloud or Redshift on AWS, with dbt for transformations | Pay-per-query pricing can surprise you. Partitioning and query limits matter from day one. |
| Lakehouse | Large raw event data, files and logs, ML training sets, keeping full history cheaply | Cloud Storage or S3 with an open table format such as Apache Iceberg or Delta Lake, queried by a warehouse or Spark | More moving parts. Only worth it when volume or variety justifies the extra operations. |
How do you stop bad data reaching reports and models?
Test every load before it's published. Check row counts, nulls, duplicates, value ranges and freshness, block the load or alert someone when a check fails, and agree in writing with source-system owners which fields they can't change without notice.
Most data incidents we see aren't exotic. A source app adds a status value, a CSV export changes its date format, or a job silently loads zero rows on a public holiday. Tests catch these in minutes instead of weeks. We put them in the same repository as the transformations, so they're reviewed together.
Data contracts sound formal but can be a short document: the table, the fields, their types and meanings, who owns them, and how changes are announced. The point is that the people who change source systems know someone downstream depends on them.
Moving off spreadsheets and legacy databases
Spreadsheets aren't the problem. Spreadsheets as the system of record are. We map which files people really rely on, design a structured model behind them, build forms or imports so data is entered once, and run old and new side by side until the totals reconcile.
Legacy databases get the same treatment: profile the data, document what the columns really mean (rarely what the names say), migrate in stages and keep a rollback path.
Keeping the cloud bill under control
- Partition and cluster large tables so queries scan only what they need.
- Load incrementally instead of rebuilding everything nightly.
- Set query quotas and budget alerts per project or team.
- Move cold data to cheaper storage tiers, and delete what nobody needs.
- Review the top ten most expensive queries every month. Usually two or three of them account for most of the spend.
Making data usable for AI
Models fail quietly when the data they see in production differs from the data they were trained on. Data engineering for AI is mostly about removing that gap.
For classical machine learning, that means features computed by the same code in training and serving, versioned so you can reproduce last quarter's model. A feature store can help when several models share features. For a single model, a well-modelled warehouse table is often enough.
For LLM applications using retrieval-augmented generation, the source documents need cleaning, sensible chunking, metadata for filtering and permissions, and a vector store kept in sync with the source. We often use pgvector inside Postgres or Supabase, because it keeps permissions and data in one place. Dedicated vector databases make sense at larger scale. See GenAI solutions for the application side and AI/ML engineering for model work. If you need labelled training data, our data labelling service plugs into the same pipelines.
Audit first, then build in thin slices
Discovery call
A free 30-minute call, then a scoped proposal within two working days.
Data audit
One to three weeks mapping sources, volumes, owners, current reports and known pain. You get a target architecture and a costed plan, sized for where you are now.
First slice end to end
One source, one model, one report or feature set, fully tested and orchestrated. It proves the design before we copy the pattern across.
Expand and migrate
Add sources and retire spreadsheets or legacy tables in stages, running old and new in parallel until the numbers match.
Hand over or run
Documentation, a data catalogue and runbooks for your team, or we stay on as an ongoing team.
How do engagements work, and how fast can you start?
Data platforms tend to start with a scoped audit and first slice, then move to a small ongoing team as more sources come in.
| Model | Team size | Start time | Notice | Good for |
|---|---|---|---|---|
| Dedicated team | 3–15+ engineers | 2–4 weeks | 30 days' written notice | Building and running a data platform over many months, with a steady stream of new sources |
| Staff augmentation | 1–5 engineers | 5–10 working days | 30 days' written notice | Adding a data engineer to a team that already owns its platform |
| Project | Scoped to the work | 1–3 week discovery | Ends on milestone acceptance | A defined migration or a first warehouse with a fixed set of sources |
| PoC or MVP | Small squad | 4–8 week build | Ends on milestone acceptance | Proving one pipeline and dashboard, or a RAG data layer, before a bigger build |
How pricing works. We quote after a free 30-minute discovery call and send a written proposal within two working days. The price depends on the seniority mix, team size, how long the engagement runs, the stack, and any compliance scope. Full details on engagement models.
When we're not the right fit
If you need a petabyte-scale platform with sub-second streaming across dozens of regions, look for a specialist with that track record; we haven't published work at that scale. If one off-the-shelf connector tool and a BI licence would solve your problem, we'll say so and help you set it up in a few days rather than sell you a custom platform. And if nobody on your side can own the data definitions, any platform we build will drift, so that conversation comes first.
Related work
Real-Time Monitoring System (RTMS)
The Town Development Fund tracked municipal infrastructure projects through manual reports, spreadsheets and field visits. We built RTMS, a web platform on PostgreSQL that replaced those reports with a standardised data model and digital forms, aggregates periodic updates into annual progress views, and gives TDF headquarters, regional offices and municipal staff role-specific access to the data relevant to them.
Questions about data engineering
What do data engineering services cover?
Getting data out of the systems where it is created, cleaning and modelling it, and storing it somewhere people and models can query it reliably. That includes batch and streaming pipelines, warehouse or lakehouse design, data quality checks, orchestration, access control, migration off spreadsheets, and keeping the cloud bill under control.
Do we need a data warehouse, or is Postgres enough?
For many companies under a few hundred gigabytes of analytical data, a well-indexed Postgres or Supabase database with a separate reporting schema is enough and costs far less. Move to BigQuery or Redshift when analytical queries start slowing the application, or when data volumes and the number of analysts grow.
Batch or streaming: which do we need?
Batch, most of the time. Hourly or daily loads cover reporting, forecasting and most ML training. Streaming is worth its extra cost and operational effort when a decision has to happen within seconds, such as fraud checks, live operational dashboards or alerting. We often run both, with streaming only for the few feeds that need it.
Can you migrate us off spreadsheets without stopping the business?
Yes. We map the spreadsheets that people actually use, build a structured data model and input forms alongside them, run both in parallel for a reporting cycle, reconcile the totals, and then retire the spreadsheets. Nobody is asked to switch until the new numbers match the old ones.
How do you make data ready for AI and RAG?
For classical ML that means consistent, versioned features that are identical in training and production. For retrieval-augmented generation it means clean source documents, sensible chunking, metadata for filtering and access control, and embeddings in a vector store such as pgvector, refreshed when the source changes.
Which tools and clouds do you work with?
We work mostly in Python and SQL, on Postgres and Supabase, and on Google Cloud and AWS, using BigQuery, Redshift, Cloud Storage or S3 as the project needs. For transformation and orchestration we commonly use dbt, Airflow or Dagster. We choose based on what your team can maintain after we leave.
Not sure where your numbers come from?
Tell us which reports people argue about and where the data lives. In 30 minutes we'll sketch what a cleaner setup would look like and what the first slice would cost.
