The whole system,
working together.

Connect the business decision to the application, the data and the platform. Build useful capabilities, understand their behaviour and improve them with evidence.

How we measure quality, performance and value

AI strategy and adoption

Decide where AI can make a useful difference and what it will take to deliver it. Connect the business case to data readiness, technical choices and the people who will use the system.

All expertise areas

Opportunity and readiness assessment

Map the work, identify bottlenecks and assess the available data. Prioritise opportunities by value, feasibility, risk and the effort needed to change the process.

Technology choices and economics

Compare conventional automation, predictive models and generative AI. Assess build-versus-buy options, provider dependencies and total cost, including data preparation, integration, evaluation and ongoing operation.

Roadmaps and organisational adoption

Define a focused first release and acceptance criteria. Redesign the workflow with the people using it, provide role-specific training and track adoption, feedback and the support needed to sustain the change.

Operating model and value tracking

Agree who owns the product, data, risk decisions and support. Set a baseline, assign responsibility for benefits and review outcomes against the full cost of delivery. Use that evidence to continue, adjust or stop an initiative.

Software and product engineering

Design, build and modernise web applications, APIs and business systems. Bring product design, sound architecture and disciplined delivery together, whether the work involves AI or conventional software.

All expertise areas

Product, service and interaction design

Map user needs and the full service journey, then test assumptions with prototypes and usability research. Develop accessible interfaces and reusable design systems. Define content, navigation, keyboard behaviour and loading, empty and error states alongside the implementation.

Application architecture and modernisation

Define domain boundaries, data models and API contracts around the business. Choose the simplest architecture that meets the requirements, document tradeoffs and evolve it as the product grows. Modernise incrementally, with compatibility checks and rehearsed data migrations.

Code quality and collaborative development

Keep changes small and reviewable, with clear naming, consistent conventions and purposeful refactoring. Use peer review, static analysis and dependency checks to catch problems early. Build input validation, access controls and safe handling of secrets into everyday development.

Test-first development and quality assurance

Use test-driven development (TDD) to define behaviour before implementation. Combine focused unit tests with integration and contract tests, then exercise critical journeys end to end. Check accessibility, performance and security against agreed requirements; reproduce defects in tests before fixing them.

Continuous delivery and production care

Automate build, test and release checks in CI/CD. Use reproducible environments, infrastructure as code and staged releases with a recovery plan. Agree reliability targets, monitor errors and latency, and maintain runbooks, dependencies and documentation so the team can operate the system after handover.

AI-assisted software delivery

Introduce coding agents and assistants across discovery, implementation, testing and maintenance. Set repository context, permissions and review standards, then assess delivery time, defects and rework against a baseline.

Agents and workflow engineering

Turn a business process into a system that can reason, use tools and complete useful work. Make the boundaries clear: what the agent knows, what it can change and when it needs a person.

All expertise areas

Orchestration and state

Design single-agent and multi-agent workflows with persistent state, scoped memory and checkpoints. Set execution budgets, retries and recovery paths so long-running tasks can resume safely.

Tool use and agent coordination

Connect tools through APIs and the Model Context Protocol (MCP), and independent agents through Agent2Agent (A2A) where useful. Validate inputs, scope permissions and define explicit delegation contracts.

Context, memory and reusable skills

Separate working context from persistent memory, summarise long sessions and load relevant instructions on demand. Package repeatable procedures as versioned agent skills, with clear scope and review.

Human oversight and controlled execution

Build approvals and intervention paths into consequential actions. Where APIs are unavailable, assess browser or computer-use agents in isolated environments with restricted access and recovery controls.

RAG, knowledge graphs and search

Connect answers to the right evidence. Combine hybrid search, knowledge graphs and agentic retrieval according to the questions people ask, the shape of the data and the cost of maintaining it.

All expertise areas

Knowledge pipelines

Prepare documents through ingestion, parsing, chunking and embeddings. Track provenance, access permissions, updates and deletions so retrieved information stays useful and current.

Hybrid and contextual retrieval

Combine keyword and vector search, metadata filters and reranking. Add document context to chunks where it improves retrieval, and test query rewriting and relevance against representative questions.

GraphRAG and knowledge engineering

Model entities and relationships to connect evidence across sources. Use graph traversal and community summaries for relationship-heavy or collection-wide questions, with entity resolution, provenance and a plan to maintain the graph.

Agentic retrieval and context assembly

Let the system decompose questions, choose sources and retrieve again when evidence is missing. Set limits on search steps and context size, preserve citations and abstain when the available evidence is insufficient.

Structured data and analytical questions

Connect natural-language questions to approved database views through constrained text-to-SQL. Validate queries and metric definitions, enforce access checks and keep the resulting evidence available for review.

Choose retrieval around the question.

These approaches can work together. We compare them on your questions and source data, including answer quality, response time and the cost of keeping the system current.

Hybrid search and reranking
Useful whenFinding specific facts and passages across documents, including exact terms and semantic matches.
What to assessRetrieval relevance, chunk context and ranking quality. A useful baseline for comparing more complex approaches.
GraphRAG
Useful whenConnecting entities across sources, following relationships or summarising themes across a collection.
What to assessGraph quality, extraction and indexing cost, access rules and the effort to keep relationships current.
Agentic retrieval
Useful whenQuestions that need several searches, multiple sources or follow-up investigation before an answer is possible.
What to assessSearch-step budgets, latency, evidence coverage and whether the system knows when to stop.
Long-context approaches
Useful whenWorking over a bounded set of relevant documents that can fit within the model’s usable context.
What to assessAttention to detail, context limits, token cost and freshness. Compare against retrieval on the same tasks.

Model selection and applied machine learning

Choose the model and level of adaptation that the problem needs. Compare language models, predictive methods and simpler baselines against your data, constraints and definition of success.

All expertise areas

Model selection and prompting

Benchmark hosted and open-weight models on representative tasks. Develop versioned prompts and structured outputs; use evaluation-led prompt optimisation and set reasoning budgets against quality, latency and cost.

Fine-tuning and model adaptation

Assess supervised fine-tuning, LoRA and preference optimisation when curated data justifies adaptation. Evaluate distillation and quantisation for smaller deployments, with held-out quality checks, licensing review and data provenance.

Prediction, ranking and decision support

Build classification, forecasting, recommendation and anomaly-detection models. Use appropriate baselines, separate training and evaluation data, and account for uncertainty and the cost of a wrong prediction.

Document, vision and voice intelligence

Work with information in the form it arrives: documents, images and audio. Connect model outputs to structured data, source evidence and a review process.

All expertise areas

Parsing, extraction and classification

Combine OCR, document layout analysis and vision-capable models to classify files and extract information. Validate outputs against schemas and business rules.

Voice and audio experiences

Connect speech recognition, language understanding and speech generation. Design for response latency, interruptions, consent and handoff to a person, with evaluation across realistic recording conditions.

Data quality and review

Keep source references, measure field-level accuracy and route uncertain results to people. Curate labelled examples for evaluation and assess synthetic data before using it.

Business systems and data integration

Bring AI into the systems people already use. Connect operational data, business applications and communication channels with clear contracts, permissions and recovery paths.

All expertise areas

Enterprise applications and channels

Connect CRM, ERP, support and collaboration workflows through approved APIs. Integration options include Salesforce, HubSpot, Microsoft 365, SharePoint, Google Workspace, Slack and Teams.

Data pipelines and synchronisation

Connect SQL databases, warehouses, object storage and vector stores. Design batch and streaming pipelines, including change-data capture where needed, with lineage, freshness checks and deletion handling.

Data foundations and ownership

Design reusable data products with named owners, shared business definitions and quality expectations. Modernise warehouse or lakehouse foundations, document data contracts and make approved datasets discoverable through a catalogue.

APIs, events and MCP

Build service APIs, webhooks, queues and MCP servers. Use scoped identities and OAuth where appropriate, manage secrets, and make retries safe through idempotency and explicit error handling.

Cloud architecture and deployment

Build on AWS, Google Cloud or Microsoft Azure around your existing estate and operating requirements. Choose managed AI services, containers or private hosting according to the workload.

All expertise areas

Architecture and deployment choices

Assess managed model services, serverless applications, Kubernetes and dedicated inference. Plan hybrid or on-premises deployment where data boundaries or existing infrastructure call for it.

Infrastructure and identity

Provision repeatable environments with infrastructure as code, including Terraform. Design network boundaries, workload identities, secrets, encryption and regional placement together.

Data residency and provider choice

Map where data, model processing and logs reside, including provider retention and access arrangements. Assess private deployment, portability and an exit plan so technology choices remain compatible with your requirements.

Reliability and cloud economics

Plan scaling, capacity, recovery and service objectives. Allocate costs by product or team, set budgets and use measured demand to guide resource choices and FinOps decisions.

AWS

Managed model access, agent services and custom machine learning in your AWS environment.

Amazon Bedrock, Bedrock AgentCore and SageMaker AI.

Google Cloud

Gemini and custom ML, connected to your data and deployed through managed or container services.

Gemini Enterprise Agent Platform (the evolution of Vertex AI), Cloud Run and GKE.

Microsoft Azure

AI applications and retrieval integrated with your Azure services and enterprise identity.

Microsoft Foundry, Azure AI Search and Microsoft Entra ID.

AI platforms, routing and optimisation

Create a foundation for multiple AI capabilities without making every team solve the same operational problems. Balance quality, responsiveness and cost with evidence.

All expertise areas

Semantic routing

Use intent classification or embedding similarity to send a request to the appropriate workflow, agent or knowledge source. Define confidence thresholds and a fallback for ambiguous requests; enforce authorisation separately.

Model gateways and inference

Select models by task requirements, measured quality, latency and cost. Centralise provider access, rate limits, budgets and fallbacks. Assess batching, streaming, prefix caching and speculative decoding for the serving workload.

Semantic caching

Reuse answers to sufficiently similar requests only when context, permissions and freshness allow it. Measure incorrect reuse and invalidate stale entries. Prefix caching serves a different purpose: reusing computation for shared prompt prefixes.

MLOps and LLMOps

Version data, prompts, models and configuration. Connect registries and evaluation to CI/CD, staged rollouts and rollback. Define monitoring and retraining triggers so changes remain repeatable and reviewable.

Evaluation and experimentation

Establish what good looks like for your application, then make changes against a repeatable baseline. Evaluate individual components and the complete user journey.

All expertise areas

LLM-as-judge and human calibration

Use model-based scoring with explicit rubrics, reference examples and human-labelled cases. Check judge agreement and review disagreements before relying on the scores.

Datasets and regression checks

Build evaluation sets from representative tasks, edge cases and observed failures. Keep held-out examples and run checks when prompts, models or retrieval change.

Agent simulations and outcome checks

Test multi-turn tasks in controlled environments, inspect tool use and verify the final system state. Repeat trials to understand variability and check recovery, permissions and escalation under failure.

Experiments and release decisions

Compare configurations against acceptance criteria. Use shadow traffic or controlled A/B experiments where appropriate, alongside offline evaluation, sampled production review and user feedback.

Observability and meaningful metrics

Make behaviour visible from the first request to the final outcome. Give product and engineering teams a shared view of quality, performance and the reasons behind failures.

All expertise areas

Tracing across the system

Connect model calls, retrieval, tool execution and agent steps in a trace, with version information and appropriate redaction of sensitive data.

Quality and operational monitoring

Build dashboards and alerts around task outcomes, evaluation results, latency distributions, errors, token usage and cost.

Feedback that improves the product

Monitor changes in input data and model behaviour, segment results by use case and investigate drift. Feed reviewed failures and user feedback into evaluation datasets and retraining decisions.

Security, responsible AI and governance

Translate the organisation’s requirements into controls people can understand and operate. Define ownership, permitted behaviour and what happens when something goes wrong.

All expertise areas

Permissions and data boundaries

Apply access controls to knowledge, tools and model interactions. Design tenant isolation, data retention, redaction and secret handling around the information each workflow actually needs.

Guardrails and adversarial testing

Red-team prompt injection, data exfiltration, memory poisoning and unsafe tool use. Combine input and output validation with isolated execution, approved tools and dependencies, and clear escalation paths.

Auditability and ownership

Maintain an inventory of models and uses, record decisions and approvals, and define responsibility for releases and incidents. Support the organisation’s risk reviews with traceable evidence and documentation.

Fairness, explainability and human review

Assess performance across relevant user groups and document limitations and uncertainty. Provide explanations appropriate to the decision, and give people a practical way to question results or request human review.

Metrics that answer
useful questions.

Define success for the task, establish a baseline and choose measures that help the team make decisions. These are examples to select from for each application.

Answer and retrieval quality

Is the response useful and supported by the evidence?

Example measures
Answer relevance, faithfulness to sources, citation accuracy, retrieval precision and recall.

Use appropriate reference data and human review. An automated judge’s score is an assessment to validate, not ground truth.

Agent effectiveness

Does the workflow complete the intended task?

Example measures
Verified task completion, tool-call correctness, repeat-run consistency, human escalation and approval outcomes.

Interpret results in context: a timely handoff to a person can be the correct outcome.

Routing and cache correctness

Does each request reach the right destination?

Example measures
Intent classification accuracy, fallback frequency, eligible cache-hit rate and incorrect answer reuse.

Check ambiguous requests, stale context and permission boundaries. A high hit rate alone does not establish correct behaviour.

Predictive performance

Are predictions useful for the decision being made?

Example measures
Precision and recall, probability calibration, forecast error and ranking quality.

Choose measures for the cost of mistakes. Account for class imbalance, time-based validation, data leakage and drift.

Speed and reliability

Does the system respond reliably when people need it?

Example measures
Time to first token, end-to-end p50 and p95 latency, throughput, timeouts and error rates.

Look at distributions and individual workflow stages. An average can hide a poor experience for some users.

Cost and product value

Is the system delivering value at a sustainable cost?

Example measures
Cost per completed task, total operating cost, adoption, time saved after review and correction, and progress against the agreed business outcome.

Assess savings alongside quality and task outcomes. A cheaper response is only useful if it still does the job.