Model intelligence inside the product you already run
Add model capability to existing products and knowledge environments with retrieval, permissions and grounded context. The systems and controls you already have stay in place; the intelligence goes where the work is, not into a chat box bolted to the side.
What the integration involves
The model is the smallest decision in this list. Everything else is what determines whether the feature is reliable, affordable and defensible.
Retrieval over your own content
Documents, records, tickets, policies and the knowledge that currently lives in three systems and one person’s head. Chunking, indexing and ranking designed around how the content is actually written.
Permissions and entitlements
Retrieval that respects who is allowed to see what. Permissions are modelled first and the index is built second, because a model that can reach everything is a data-protection incident waiting for a prompt.
Grounded answers with citations
Output tied to the source passages it came from, so a reader can check it. Ungrounded fluency is the single most expensive failure mode in this category.
Context and prompt engineering
What goes into the window, in what order, at what cost, with the assembly written as code that can be tested rather than a string somebody edits in production.
Model selection, routing and fallback
Which model handles which class of request, what happens when a provider degrades, and how you move between them without rewriting the product.
Cost and latency budgets
Tokens per unit of work, caching, batching and an agreed response-time target, so the feature is affordable at the volume you actually expect.
All of them are design problems, not model problems
Every one of these is recoverable, and every one is considerably cheaper to prevent than to discover in front of a customer.
It answers confidently from nothing
No grounding, no citation, no refusal path. The fix is retrieval design and an explicit “I do not know”, not a sterner prompt.It answers from documents the user may not see
Entitlements were applied to the application and not to the index. This is the failure that stops enterprise rollouts.It was right in the demonstration
A handful of questions is not a test set. Without a graded golden set you have an anecdote with a launch date.It cost four times the estimate
Context assembled without a budget, no caching, and a model chosen for the hardest request and used for all of them.An answer nobody can check is an opinion with a citation style
The purpose of retrieval is not to make the model sound informed. It is to make the answer traceable: this passage, from this document, at this version, at this time. That is what lets a professional rely on the output, and it is what lets you explain the output later to somebody who is unhappy about it.
It also makes the system measurable. Grounded answers can be graded pass or fail against an agreed definition of correct, which is what turns a feature into something you can hold to a threshold.
Most of the difficulty in an LLM integration is upstream of the model: the content has to be reachable, the permissions have to be modelled, and there has to be a stable vocabulary between your systems and the retrieval layer so an answer does not change because somebody renamed a field.
Where that layer does not exist yet, we build it, and we are explicit that it is the larger part of the work rather than presenting it as a preliminary.
Questions worth answering
What is RAG and when is it the right approach?
Retrieval-augmented generation means the model answers from passages retrieved out of your own content at the time of the question, rather than from what it absorbed in training. It is the right approach when answers must reflect current, organisation-specific material and must be checkable against a source, which covers most enterprise knowledge use cases.
Can you add AI to a product we already have?
Yes, and it is one of the most common engagements. Model or agent capability is integrated into an existing application, knowledge base or workflow while preserving the systems, permissions and controls already in place, rather than requiring the product to be rebuilt around the feature.
Can you build a chatbot for our business?
Usually what is wanted is not a chatbot in the 2017 sense but an assistant that answers from your own documents, records and policies, respects who is allowed to see what, and cites the passage it answered from. That is retrieval work rather than conversation design, and it is what this page describes. Where the need genuinely is a scripted flow over a handful of known questions, a simpler tool is cheaper and we will say so rather than build a model into it.
How do you stop a model answering from documents a user is not allowed to see?
By modelling permissions before the index is built, so entitlements are enforced at retrieval rather than applied afterwards in the application layer. Retrieval that respects who is allowed to see what is the part of enterprise RAG that most often decides whether a rollout goes ahead.
Have a product that should be answering questions?
Tell us what it holds, who is allowed to see which parts of it, and what a good answer looks like. Those three answers decide the architecture.