Models you can explain to an auditor

Our AI/ML engineering services start with a business question and a measured baseline, then build, evaluate and deploy the model that beats it. Every model ships with an evaluation report, a model card and a runbook your team can use.

  • Framing and baseline in 1–3 weeks
  • Evaluation report on held-out data
  • Model card and retraining runbook
What we do

From a question to a model in production

Most failed ML projects fail before any training starts: the question was vague, nobody measured the simple alternative, or the data couldn't support the answer. So we spend the first weeks on framing and baselines, and only then pick a model.

Problem framing and baselines

We turn "use AI on our churn data" into a decision someone makes, the metric that matters for it, and a baseline to beat. Often the baseline is a rule or a logistic regression. Sometimes it's good enough and you don't need us for the rest.

Data preparation and features

Cleaning, joining and labelling the data, finding leakage before it inflates your results, and building features that will still exist at prediction time. If the pipelines need work first, our data engineering team handles it.

Model selection

Classical ML, deep learning or an LLM, chosen on cost, latency, explainability and how much data you have. We show you the trade-off in numbers rather than defaulting to the newest thing.

Training and evaluation

Held-out test sets that stay untouched until the end, time-based splits for anything that changes over time, and error analysis by segment so you know where the model is weak before your users find out.

Deployment and serving

Batch scoring, a real-time API or an on-device build, packaged in Docker and deployed to your Google Cloud or AWS account. We load-test it and agree latency and cost targets before go-live.

Handover and documentation

Training code, evaluation notebooks, a model card that states known limits, and a runbook for retraining. Your engineers should be able to rebuild the model without calling us.

Should you use classical ML, deep learning or an LLM?

Use classical ML for tabular data and clear numeric targets, deep learning for images, audio and signals where you have plenty of labelled examples, and LLMs for unstructured text and language tasks. Start with the cheapest option that could plausibly work and only move up when the evaluation shows a real gain.

ApproachUsually the right call whenWatch out forTools we commonly use
Classical MLData is rows and columns: transactions, sensor readings, CRM fields. You need forecasts, scores or classes, and you need to explain them.Feature work takes most of the effort. Weak when the signal lives in free text or images.scikit-learn, XGBoost or LightGBM, pandas
Deep learningImages, audio, video, sensor sequences, or text where you have enough labelled examples to fine-tune a model you control.Needs more labelled data and GPU time. Harder to explain, and serving costs more.PyTorch, pretrained vision and audio models, ONNX for export
LLMsUnstructured documents, extraction, summarisation, search over your content, conversational interfaces.Per-call cost, latency, and wrong answers delivered confidently. Needs an evaluation set and guardrails from day one.Hosted model APIs, open-weight models, RAG pipelines (see GenAI solutions)

A pattern we often recommend is a hybrid: an LLM pulls structured fields out of messy documents, and a small classical model makes the actual decision on those fields. It's cheaper to run and easier to audit.

How do you know a model is good enough to ship?

A model is ready when it beats the baseline on a held-out test set, using a metric that reflects what a wrong prediction costs your business, and when its errors are acceptable in every segment you care about, not just on average.

Accuracy alone rarely tells you much. If you're flagging fraudulent claims, a missed fraud and a wrongly blocked customer cost different amounts, so we pick the threshold based on those costs and report precision and recall at that point. For forecasts we report error in the units your planners use, such as units of stock or rupees, not just a percentage.

We keep a final test set locked away until the end and check performance by region, product line or customer type. If the model only works for your biggest segment, you should know before launch.

Deployment patterns: batch, real-time API or edge

Pick the simplest pattern that meets the decision's deadline. If a prediction is only needed by tomorrow morning, a nightly batch job is cheaper and easier to run than an API.

PatternFitsTypical stackTrade-off
Batch scoringDaily risk scores, demand forecasts, lead ranking, report enrichmentPython job in Docker, scheduled on Google Cloud or AWS, results written to your database or warehouseCheapest and simplest. Predictions are only as fresh as the last run.
Real-time APIScoring at checkout, recommendations in an app, document processing on uploadFastAPI service in Docker on Cloud Run, ECS or Kubernetes, or a managed endpoint such as Vertex AI or SageMakerFresh answers, but you now own latency targets, scaling and uptime.
Edge or on-deviceOffline field apps, cameras, sensors, places with poor connectivityQuantised or pruned models exported to ONNX or TensorFlow LiteNo network round trip. Model size and update rollout become the hard part.

Whichever pattern you choose, the model ships with logging of inputs and outputs, a version tag, and a way to roll back. Keeping it healthy after launch is the job of MLOps and AI infrastructure.

How an engagement runs

Five stages, with a go or no-go point after the baseline

  1. Discovery call and data review

    A free 30-minute call, then a scoped proposal within two working days.

  2. Framing and baseline

    One to three weeks. We agree the decision, the metric and the baseline, and check the data can support the question. You get a short written finding and a recommendation, including "don't build this" if that's the honest answer.

  3. Build and evaluate

    Feature work, model training and error analysis in short cycles, with results shown against the baseline every week or two.

  4. Deploy

    Packaging, load testing and release to your cloud account, behind a flag or in shadow mode first where the risk justifies it.

  5. Hand over or keep running

    Documentation and a walkthrough for your team, or we stay on to run and improve the model as part of an ongoing team.

How do engagements work, and how fast can you start?

Most ML work starts as a scoped project or PoC and moves to a dedicated team once there is a model worth maintaining.

ModelTeam sizeStart timeNoticeGood for
Dedicated team3–15+ engineers2–4 weeks30 days' written noticeA roadmap of several models, or a model that needs continuous improvement
Staff augmentation1–5 engineers5–10 working days30 days' written noticeAdding one or two ML engineers to a team that already has an ML lead
ProjectScoped to the work1–3 week discoveryEnds on milestone acceptanceOne well-defined model with a clear metric and deployment target
PoC or MVPSmall squad4–8 week buildEnds on milestone acceptanceTesting whether ML can beat your current process before committing

How pricing works. We quote after a free 30-minute discovery call and send a written proposal within two working days. The price depends on the seniority mix, team size, how long the engagement runs, the stack, and any compliance scope. Full details on engagement models.

When we're not the right fit

If what you need is a dashboard or a few SQL reports, you don't need a model and we'll tell you so. If you have no historical data and no way to collect labels, start with data labelling or a manual process that generates data, and come back when there is something to learn from. And if you want a guaranteed accuracy number in the contract before anyone has seen the data, we're the wrong team: we'll commit to a measured baseline and an honest go or no-go, not a figure we made up.

Related work

Research unit

Dignep AI Research Unit (DAIR)

DAIR is our applied AI research wing, led by Dr. Yagya Raj Pandeya of Kathmandu University's Department of Artificial Intelligence. Its work includes retrieval-augmented generation for Nepali and regional languages, multimodal models, and edge AI using quantisation and pruning.

About DAIR

Case study

Certifyi AI SaaS GRC platform

We built Certifyi's compliance automation platform end to end. Its stack includes Python for the AI and automation layer, with TensorFlow or PyTorch where applicable, running on AWS with Docker and CI/CD through GitHub Actions.

Read the Certifyi case study

Questions about AI/ML engineering

What do AI/ML engineering services include?

They cover the work between a business question and a model running in production: framing the problem, building a baseline, preparing data, choosing and training a model, evaluating it on data it has never seen, deploying it as a batch job, API or edge build, and handing over code and documentation your team can run.

Do we need an LLM, or will a classical model do?

If your inputs are rows and columns and you want a number or a category out, a classical model such as gradient-boosted trees is usually cheaper, faster and easier to explain. LLMs earn their cost on unstructured text, documents and conversation. We test the cheaper option first and show you the comparison.

How much data do we need to start?

There is no fixed number. It depends on how many outcomes you are predicting and how noisy the labels are. In the first one to three weeks we look at what you have, build a baseline and tell you whether more data, better labels or a different approach would move the result most.

Can you guarantee a model accuracy figure up front?

No, and we would be wary of anyone who does before seeing your data. What we commit to is a measured baseline, an agreed metric tied to the business decision, and a clear go or no-go point once we have evaluated a first model on held-out data.

Who owns the models and code?

You do. Code goes into your repository, models and data stay in your cloud account where possible, and the handover includes training scripts, evaluation notebooks, a model card and a runbook. We do not lock models behind a hosted service of ours.

How is this different from MLOps?

AI/ML engineering gets a model built, evaluated and into production once. MLOps keeps it working: automated retraining, model versioning, monitoring for drift and cost, and safe rollouts of new versions. Many clients start with engineering and add MLOps once the first model has proved its value.

Have a question you think a model could answer?

Send us the question and a description of the data you have. In a 30-minute call we'll tell you whether ML is likely to help, what a baseline would look like, and what the first few weeks would cost.

Scroll to Top