Skip to main content
NELLA Labs

Services · Automate

AI systems with evaluation, governance and a human in the right place.

The hard part of applied AI is not the model. It is knowing which problem is worth automating, grounding the system in your actual data, measuring whether the output is good, and designing what happens when it is wrong. We build AI that is evaluated, observable and governed — and we will tell you when the answer is not AI at all.

Problems we hear

If you recognise any of these, this is the right page

  • We are under pressure to “do something with AI” but cannot separate the real opportunities from the noise.
  • We built a prototype that demonstrated well and then could not be trusted in production.
  • Our staff spend hours on document handling, triage and re-keying between systems.
  • We cannot tell whether our AI feature is getting better or worse over time.
  • Our legal and risk teams will not approve an AI system we cannot explain or audit.

Outcomes

What you should have at the end

A ranked opportunity list
Use cases scored on value, feasibility, data readiness and risk — including the ones we recommend you do not pursue.
Evaluated quality
An evaluation suite that runs in CI, so a prompt or model change either improves measured quality or fails the build.
Governed operation
Recorded model, prompt version, inputs, outputs, cost, latency and human overrides for every AI run.
Measured time recovered
Automation instrumented against a before-and-after baseline, so the business case can be verified rather than asserted.

Capabilities

What Automate covers

Each of these is a distinct piece of work with its own deliverables. They combine into engagements rather than being sold separately.

01

AI opportunity and readiness assessment

Find the use cases worth funding, and prove your data can support them.

The problem

AI programmes commonly start from the technology rather than a costed problem, and stall when the underlying data turns out to be unusable.

What the work involves

We run workshops with the people doing the work, quantify the current cost of each candidate process, and assess data availability, quality, ownership and lawful basis for each. Every use case gets a score across value, feasibility, data readiness and risk, and a recommendation — including an explicit "do not do this" where that is the honest answer.

Typical deliverables

  • Scored and ranked use-case register
  • Data readiness assessment per use case
  • Risk and regulatory screening, including DPIA triage
  • A sequenced roadmap with a recommended first proof of value
02

Generative AI applications

Production applications with grounding, guardrails and a defined failure mode.

The problem

Demonstrations succeed on curated inputs. Production fails on the long tail, and an ungrounded model will answer confidently rather than admit it does not know.

What the work involves

We ground generation in your own content, constrain outputs to structured schemas where the downstream system needs reliability, and design explicitly for refusal and escalation. Provider access sits behind a gateway abstraction so a model change is a configuration decision, and cost and latency are controlled at the boundary.

Typical deliverables

  • Model gateway with provider abstraction and cost controls
  • Prompt and schema versioning with structured output validation
  • Guardrails, redaction and refusal behaviour
  • Human review paths for consequential output
03

Enterprise search and RAG

Answers grounded in your documents, with citations and respect for permissions.

The problem

Naive retrieval returns plausible answers from documents the asker should not be able to read, and cites nothing you can check.

What the work involves

We build retrieval that enforces the source system’s permissions at query time rather than at indexing time, so an answer can never be assembled from content the user cannot access. Every answer carries citations to its sources. Retrieval quality is measured against a labelled evaluation set rather than assessed by impression.

Typical deliverables

  • Ingestion pipeline with permission-aware indexing
  • Hybrid retrieval with reranking, tuned against a labelled set
  • Cited answers with source-document links
  • Retrieval and answer-quality evaluation suites
04

AI agents and copilots

Assistants that take real actions, within boundaries you set.

The problem

An agent that can act on your systems is a security and safety design problem long before it is a model problem.

What the work involves

We define the tool surface narrowly and explicitly, run every action through the same authorisation checks a human user would face, and require confirmation for anything consequential or irreversible. Every action is logged with its inputs, the reasoning available, and the identity it acted on behalf of.

Typical deliverables

  • Scoped tool definitions with least-privilege authorisation
  • Confirmation and approval gates for consequential actions
  • Full action audit trail with attribution
  • Escalation to a human with the full context preserved
05

Voice and conversational AI

Conversational interfaces on the channels your users already use.

The problem

Chatbots deployed without a handover path trap users in loops and damage trust more than having no bot at all.

What the work involves

We design the handover to a human first and the automation second. Conversations carry their full context across the handover. AI participation is disclosed to the user. We support web chat, WhatsApp, SMS and voice through provider interfaces, and we build against mock adapters where credentials are not yet available rather than pretending an integration exists.

Typical deliverables

  • Conversation design with explicit escalation paths
  • Channel integrations behind provider interfaces
  • AI disclosure and consent handling
  • Transcript retention, moderation and reporting
06

Document intelligence

Extract structured data from documents, with a confidence threshold and a review queue.

The problem

Extraction accuracy is never one hundred per cent, and a pipeline with no review step quietly writes wrong data into systems of record.

What the work involves

We define per-field confidence thresholds and route anything below them to a human review queue with the source document visible alongside the extracted values. Corrections feed the evaluation set. Accuracy is reported per field and per document type rather than as a single headline number.

Typical deliverables

  • Classification and extraction pipeline
  • Confidence thresholds with a human review queue
  • Per-field accuracy reporting
  • Retention and secure deletion for source documents
07

Workflow and process automation

Remove the re-keying, the chasing and the spreadsheet in the middle.

The problem

Significant staff time goes on moving data between systems that were never integrated, and on chasing approvals through email.

What the work involves

We map the process as it is actually performed, including the workarounds, then automate the mechanical parts and leave judgement with people. Automations are durable: they record trigger, version, inputs, result, error, retries and any human override, and they respect communication consent and quiet hours.

Typical deliverables

  • Process map of the as-is and to-be flows
  • Durable, observable automation with retry and dead-letter handling
  • Approval and exception routing
  • Before-and-after measurement against a baseline
08

AI evaluation, observability and governance

Know whether your AI is working, and be able to prove it.

The problem

Without evaluation, AI quality is a matter of anecdote, and a prompt change made on a Friday can silently degrade output for months.

What the work involves

We build labelled evaluation sets from your real cases, run them in CI on every prompt or model change, and track quality, cost and latency over time. Every run is recorded with model, prompt version, inputs, outputs, cost, latency, feedback and human overrides — which is also what your risk function needs to approve the system.

Typical deliverables

  • Labelled evaluation sets and automated scoring in CI
  • Per-run records: model, prompt version, cost, latency, outcome
  • Quality, cost and drift dashboards
  • AI register, governance documentation and review workflow

Engagement models

How we can work together

AI use-case workshop

A fixed-fee workshop producing a scored use-case register and a recommended first step.

Best for: Organisations that need to separate real opportunities from pressure to act.

Proof of value

A time-boxed engagement that builds one use case far enough to measure it against a defined success threshold — and stops if the threshold is not met.

Best for: Proving a business case before committing to a production programme.

Production programme

Milestone-based delivery of a production AI system with evaluation and governance included.

Best for: A validated use case moving into real operational use.

Outcome-linked models

Where the outcome is genuinely measurable and attributable, part of the fee can be linked to it. We only propose this where the measurement is not gameable.

Best for: Automation with a clean, agreed baseline.

Delivery process

How the work runs

  1. 01

    Explore

    Identify candidate use cases and the cost of the status quo.

  2. 02

    Shape

    Score use cases, assess data, screen risk, define success thresholds.

  3. 03

    Build

    Build the system with evaluation running from the first increment.

  4. 04

    Validate

    Evaluate against the threshold; report honestly if it is not met.

  5. 05

    Launch

    Staged rollout with human review and monitoring in place.

  6. 06

    Operate

    Monitor quality, cost and drift; retrain or revise as evidence requires.

Technology approach

What we build with, and why

Model access
A provider-agnostic gateway. We have no commercial incentive to prefer one model vendor, and your data-use constraints determine what is permissible before capability does.
Retrieval
Postgres with pgvector for most workloads; dedicated vector stores only where scale genuinely requires them.
Orchestration
Durable background execution with idempotency, retry and dead-letter handling. AI work never runs in a request handler.
Evaluation
Labelled sets, automated scoring, regression gates in CI, and per-run cost and latency tracking.

Security and quality

Non-negotiables

  • AI never makes a binding commercial commitment, final price, or irreversible decision without a human.
  • AI-generated content is disclosed as such and is always editable by the person it is presented to.
  • Sensitive data is redacted before it leaves our boundary, and data-use constraints are configured per provider.
  • Every AI run is recorded so a decision can be reconstructed months later.
  • We build evaluations before we expand a high-impact use case, not afterwards.

Investment

Starting points

Indicative bands, not quotations. What moves a project within — or outside — these ranges is scope, integration count, data quality and regulatory context.

AI use-case workshop

£2,500 – £6,000

Scored register and a recommended first step.

AI proof of value

£15,000 – £40,000

One use case built and measured against a defined threshold.

Production AI assistant / RAG

£40,000 – £120,000

Grounded, evaluated, governed and integrated.

Multi-agent / workflow platform

£80,000 – £250,000+

Multiple automated processes with orchestration and governance.

All published figures are indicative and exclude tax, cloud and model usage, third-party licences and app-store fees. A price becomes an offer only when a person at NELLA Labs confirms it in writing.

Related

Where this shows up

NELLA Labs product

Daju Verify

Onboard people and businesses with evidence you can audit.

NELLA Labs product

Trellis

Learning that connects families, learners and verified educators.

Industry

Financial services and fintech

Onboarding, money movement and controls that survive an audit.

Industry

Government and public services

Services that must work for everyone, and be defensible afterwards.

Industry

Professional services

Utilisation, engagement records and knowledge that leaves with people.

Frequently asked questions

Will our data be used to train someone else’s model?

Not without your explicit instruction. We configure providers for zero-retention or no-training processing where the provider offers it, record which provider processes what in the subprocessor register, and redact sensitive fields before they leave our boundary. Which providers are permitted for your data is a decision you make during Shape, and it constrains the architecture rather than the other way round.

How do you stop it from making things up?

Three ways, in order of effectiveness. We ground answers in your own content and require citations. We constrain outputs to a schema so downstream systems reject malformed or unsupported responses. And we design refusal and escalation as first-class behaviour, so "I do not have enough information" is an acceptable, tested outcome rather than a failure the model avoids.

Can you tell us not to use AI for something?

Yes, and we do. Some processes are better fixed with an integration, a form change or a policy decision, and some are too consequential or too poorly evidenced to automate responsibly. That recommendation is part of what a use-case workshop is for.

What does an AI system cost to run?

Model and infrastructure usage is billed separately from our fees, because it varies with your volumes. We model expected cost per transaction during Shape, instrument actual cost per run in production, and set budget alerts. You see the real number rather than an estimate that ages badly.

How do you handle regulated or high-risk use cases?

We screen every use case for high-risk processing — identity, biometrics, children, health, large-scale profiling and automated decisions with legal effect — during Shape, and flag where a DPIA is likely to be required. We build the audit trail and human-review paths that a regulator or your own risk function will ask for. We do not provide legal advice; we make sure your counsel has a system they can actually assess.

Next step

Start a automate conversation

The Project Architect will already know you came from this page, and will ask questions that fit.