Core Utilities · Architecture and operations

12 minute readReview draft

Share utilities without sharing business logic

An LLM gateway, embeddings, OCR, and transcription can remove repeated integration work. They remain useful only while the shared layer owns transport, policy, and operations—and the product retains meaning, validation, and user consequence.

The attraction

A shared endpoint can hide repetition. It can also hide responsibility.

The first product connects to a model provider. The second needs embeddings. A third needs invoices extracted through OCR and a fourth needs recorded conversations transcribed. Each integration begins as a small technical task: acquire credentials, send bytes, receive a result. Once several products repeat that work, centralising it looks economical.

The economy is real, but the request is rarely just technical. It carries data classification, retention, provider eligibility, regional constraints, cost, latency, and a product decision about what happens when the result is uncertain. If the utility absorbs all of those decisions, it becomes a hidden business platform. If it absorbs none of them, it is merely another proxy to operate.

The useful shared boundary removes repeated mechanics while making product responsibility more visible.

Thin shared layer

Centralise transport and policy. Keep meaning with the product.

Product owns

Prompt, retrieval design, user decision, domain validation, and fallback experience.

Shared utility owns

Authentication, request envelope, provider policy, quotas, audit, metering, and operational signals.

Provider owns

Model or processing runtime, provider availability, and provider-specific behaviour.

If changing a product prompt requires a utility release, the shared boundary has become too thick.

The boundary is intentionally narrow. Products express intent and evaluate meaning; the utility enforces common transport and policy; providers execute the specialised workload.

The contract

Every request should carry enough context to be governed

A generic payload is not enough. The utility needs a stable request envelope that states which product is asking, which capability it needs, the data classification, an accountable use case, latency expectations, and a correlation identifier. For generative work it may also include an approved model class and budget ceiling. For OCR or transcription it may include language, region, retention, and whether human review is required.

Centralising model routing inside an AI gateway lets the shared layer make deterministic decisions before provider execution. It can reject a restricted document for an ineligible provider, enforce product quotas, select an approved regional route, and create an audit event without attempting to understand whether the output is a valid invoice, a safe answer, or an accurate clinical term.

ConcernShared utility decidesProduct decides
Provider routeWhich approved provider satisfies policy and availabilityWhich capability and assurance the use case requires
DataTransport, encryption, retention instruction, and permitted regionClassification, minimisation, lawful use, and user disclosure
QualitySchema integrity, transport success, and provider metadataWhether the result is correct enough for the domain decision
CostMetering, quotas, provider price signals, and budget refusalValue threshold, user limits, and acceptable degraded behaviour
FailureTyped refusal or provider failure with correlation evidenceRetry experience, manual path, and whether the workflow may continue

Failure is part of the interface

A shared success path needs several different stop paths

01 · refused

Policy refusal

The request violates a data, provider, or workload rule. Return a reason, not a generic error.

02 · decision

Capacity limit

The utility is healthy but budget or concurrency is exhausted. Queue, degrade, or defer deliberately.

03 · decision

Provider failure

Retry only where idempotency and latency permit. Route only where policy still holds.

04 · refused

Invalid result

Transport succeeded, but product validation rejects the output. The product owns the fallback.

These outcomes need different product behaviour. A single 500 response collapses governance, capacity, provider reliability, and domain validation into one unhelpful failure.

Failure design

A successful call can still be an unusable result

Provider availability is the easiest condition to observe. More important failures can arrive with a successful status code: an OCR result has lost a decimal separator, a transcript attributes a sentence to the wrong speaker, an embedding model changes the geometry of an existing index, or an LLM returns fluent text that does not satisfy the product's evidence requirements.

The utility should preserve the information required to investigate those outcomes—provider, model or engine version, timings, policy route, usage, and a correlation identifier without the raw payload. This evidence trail is what makes answer playback possible: separating what was recorded at production runtime from later replays so contested outputs can be inspected rather than reconstructed. It should not declare domain correctness. The consuming product validates structured fields, confidence thresholds, citations, workflow invariants, and any action a user or system might take from the result.

  • Return typed refusals so a product can distinguish policy, capacity, provider, and validation paths.
  • Require idempotency for operations that may be retried or delivered asynchronously.
  • Version material behaviour such as model families, extraction schemas, and embedding dimensions.
  • Keep raw sensitive payloads out of ordinary logs; preserve correlation and policy evidence instead.
  • Design a manual or deterministic path for workflows that must continue when the utility cannot.

Versioning

Provider replacement is a data change as well as an integration change

A provider-neutral endpoint can make two APIs look alike while their outputs remain incompatible. Changing an embedding model alters the coordinate space of every stored vector; old and new vectors may be the same length and still be meaningless together—a critical risk for retrieval architectures that must respect access control boundaries. A new OCR engine can return the same JSON schema while segmenting pages, dates, or decimal values differently. A language model can keep the same name while a provider release changes refusal, formatting, or tool-selection behaviour.

Provider routing must therefore include semantic versions owned by the utility, not only provider model names. A capability version states the behaviour a product has tested: output schema, model or engine family, processing policy, and compatibility expectations. Products select that version explicitly. The utility may move between equivalent provider deployments inside the contract, but a material output change requires a new version and product evidence.

Capability changeMigration workEvidence before adoption
Embedding modelRebuild or isolate the index; never mix spaces by accidentRetrieval evaluation against representative queries and documents
OCR engine or schemaReprocess where required and preserve the source documentField-level accuracy, rejection behaviour, and manual correction path
Transcription engineDecide whether historic transcripts remain comparableLanguage, speaker attribution, terminology, and retention checks
LLM model or routePin prompts, tools, and fallback rules to the capability versionProduct evaluation using real task shapes and defined failure criteria

Dual-running is often safer than an instant switch. The utility can send a sampled workload to the candidate route without exposing its result to users, then give the product comparable outputs and cost signals. Sensitive payloads still follow the declared data policy, and sampled results need their own retention rule. Once the product accepts the new version, the old route needs a retirement date, a consumer list, and an owner for the remaining migrations.

Operations

Meter value and risk, not only tokens and seconds

Centralisation gives the organisation a better view of usage, but a provider invoice is not an operating model. The team needs to know which product and use case created demand, which policy decisions were made, how often users encountered a degraded path, and whether a new engine changed domain quality. Cost, latency, refusal rate, invalid-result rate, and backlog age belong beside availability.

Changes should be introduced through declared capability versions. A product can test a new OCR engine against representative documents or rebuild an embedding index before traffic moves. For LLM workloads, evaluation belongs to the product because quality is inseparable from the prompt, retrieved context, user population, and consequence. The utility provides comparable telemetry and a controlled route; it does not provide a universal quality score.

Operating signalQuestion it answers
Usage by product and use caseWhere is demand growing, and does the use case have a named owner?
Policy refusal rateAre products sending workloads the approved routes cannot accept?
Cost per completed outcomeIs cheaper inference actually cheaper after retries and manual correction?
Invalid-result and fallback rateHow often does technical success fail the product's domain checks?
Version adoptionWhich products remain exposed to an old provider, schema, or model behaviour?

The counter-case

Do not centralise an experiment before its boundary is understood

A short-lived product experiment may be clearer with a direct provider integration, especially when the team is still discovering what needs to be measured and no sensitive data or irreversible action is involved. Generalising too early freezes guesses into a shared contract and forces every later use case through assumptions made for the first one.

Move a capability into the shared layer when several products repeat the same operational and policy responsibilities—not merely when they call the same vendor. The common utility should make providers more replaceable, incidents more diagnosable, and policy more consistent. If it also needs to understand every product's business rules, the boundary has crossed too far.

Centralise the obligations that repeat. Keep the meaning that differs close to the product.

Continue through the system

What an AI gateway actually does

Why centralising model access works only while prompt and business logic remain with the application.

Read the article

Why your AI system needs playback

Preserving historical state, retrieval ranking, and provenance behind generated answers.

Read the article

RAG that respects access control

Four approaches to permission-aware retrieval and why document security belongs in the index.

Read the article

The reusable unit is assurance

How common controls can improve every product without replacing product-specific evidence.

Read the article

Let's talk about your challenge

If your organization is working with complex digital systems or exploring operational AI, we are always open to a conversation.