Documents, services and language models in operation

Running AI On-Premise or Externally: Which Data Leaves the Company?

Using an AI assistant that answers questions based on company documents as an example, this article shows which components run internally or externally and which data is transferred. On-premise here means operation on infrastructure located at the company's premises.

System KnowledgeAuthor: Hannes SchubertPublished: Last updated:

However, the location of the user interface does not tell us where the language model runs: even an internally deployed application can send document passages to an external service. What matters, therefore, is the path from the question through the information used to the answer.

What is operated in a document-based AI assistant

A language model generates text based on its inputs and learned model parameters. Training determines or changes those parameters using training data. This must be distinguished from using the trained model to process inputs and produce an output; in operational use, this is referred to as inference. ISO/IEC 22989 defines inference more broadly as reasoning and separately describes the trained model and model training.[6]

If the assistant is to take company documents into account, a search component is added. It identifies relevant passages that are passed to the language model together with the question. This principle is called Retrieval-Augmented Generation, or RAG. The BSI describes it as incorporating stored information without that information necessarily having been used as training material beforehand. This does not mean that every RAG architecture is developed without training; the original approach by Lewis and colleagues also includes jointly fine-tuning components.[2][4]

Four parts can be distinguished for the assistant considered here: the application, the document repository, the search index and search, and processing by the language model. In vector-based search, document passages and queries are converted into numerical representations known as embeddings. The index holds the representations of the passages for retrieval. Generating these embeddings is a separate processing step.[3]

Three operating arrangements for the same example

The following cases are architectural examples developed for this article for the same assistant, not a comprehensive classification of all AI systems. In each case, an employee asks for the approved maintenance instruction for a particular device. What changes is where the components run and who operates them.

Comparison caseDistribution of componentsData path in the example
A · Entirely on-siteThe application, documents, index, embedding generation and language model all run within the company.No external service is envisaged for retrieval or answer generation in this example.
B · Dedicated service-provider environmentAll these components are located in an environment at the service provider intended for the company's use.Documents and queries are processed at the service provider.
C · Internal search, external language modelThe application, document repository, index and embedding generation remain internal. The language model is called externally.The question and selected document passages are transferred to the model endpoint.

In case B, the data is held outside the company's own premises at the service provider. This also applies if the company administers the environment itself. A dedicated deployment describes the intended use of the environment; which people can gain administrative access must be clarified separately.

Self-managed operation, on-premise and cloud are therefore not interchangeable terms. Self-managed operation concerns responsibility for running the system, while on-premise here concerns location. Under the NIST definition, a private cloud can be located either on or off the organisation's premises.[1]

An API, too, is first and foremost an application programming interface, not an operating location. It can connect internal or external components. Case C is a hybrid distribution of internal and external components in this article. This is not automatically a hybrid cloud in the NIST sense: there, the term refers to a composition of several distinct cloud infrastructures.[1]

Which data is transferred when an answer is requested

Before the first question, the document collection intended for retrieval is prepared. In the vector-based setup described here, its passages are processed by an encoder and their embeddings are indexed. Karpukhin and colleagues distinguish this preparatory step from the later conversion of a specific query into a search vector.[3]

In case C, this processing initially remains internal. If an external embedding service is used instead, all document passages intended for this purpose are transferred to that service during the initial index build—not just later search results. During updates, the passages whose embeddings are recalculated are transferred; a complete rebuild may again involve the entire included collection. This is a consequence of the chosen component distribution, not a statement about where the encoders in the research approach were operated.[3]

For the question about the maintenance instruction, the application first identifies relevant passages that the employee is authorised to access. The BSI describes how permissions can be taken into account during retrieval and advises against implementing a permissions and roles system solely through textual instructions to the model.[4]

In case C, the application then sends the question together with the selected passages to the external language model. The document repository remains internal, but parts of its contents nevertheless leave the company. If the query is also converted into an embedding externally, this creates another transfer path to the embedding service. This service and the service used for answer generation need not be the same.[2][3][4]

The answer is returned to the application. For the architectural example, it must also be specified whether and where the question, passages and answer are stored in a chat history or log. The visible answer is therefore only one part of the processing chain; storage and subsequent access belong in the same assessment.

Cases A and B with all components at one location; case C with internal components and an external language model.
Architectural illustration of cases A–C: in case C, the document repository, search and embedding generation remain internal. The question and selected document passages are transferred to the external language model; the answer returns to the application. An external embedding service would add another data path. Foundations: [2–4].

What must be clarified before choosing an operating arrangement

For every service involved, it should be clear which content it receives, where it processes or stores that content, and who can access it. The NIST profile for generative AI recommends recording third parties with access to organisational content. It also addresses reviewing contractual terms and the risks of unauthorised further use of data. The label “API” alone does not establish any specific rule for storing or using inputs.[5]

A separate question concerns the model version: which version processes the query, who determines when it changes, and how are changes announced and reviewed? The BSI identifies versioning as a selection criterion and notes that the extent of control can depend on the operating arrangement and contractual terms. The NIST profile recommends continuous monitoring of third-party generative systems in use.[4][5]

For selection, this means that the available controls must be assessed for the specific offering. An unannounced change must not be assumed for every model API, nor does administering the system oneself replace a controlled change procedure. Versioned knowledge models concern a different type of model; there, too, however, it must remain clear which version a review refers to.

Operating the system also requires an answer to what happens if a component fails. Whether search remains usable, a query stays pending or an alternative workflow is provided must be clarified for the specific use. When introducing workflow software, such fallback arrangements are likewise part of preparing for routine operation.

Three questions about the company's own operational workload

The data path alone does not yet establish whether an operating arrangement suits the organisation. Three questions help with an initial assessment:

  • How evenly is the assistant used—and what peak loads are expected?
  • What happens in the event of a failure, and for how long is the planned fallback arrangement viable?
  • Who maintains the language model, search index and access permissions, and reviews changes during operation?

These questions do not yet constitute a cost calculation. They identify the needs against which a specific service offering can be assessed. No blanket recommendation for the company's own servers or an external service can be derived from them.

Reviewing the specific processing chain

A convincingly worded answer can still be wrong. Under “Confabulation”, NIST also describes outputs with fabricated reasoning or source references and recommends checking the sources and citations in generated answers. The BSI additionally points out that a stored knowledge base can itself have been manipulated. RAG is therefore no guarantee of substantive correctness.[4][5]

In the maintenance example, the answer must therefore be checked against the passages actually used and their suitability for the device concerned. For decision support, retrieving information must remain distinct from deciding on the subsequent action.

A review log could bring together the query, the document passages used, the model and configuration references, and the answer. This is an architectural proposal for subsequent review, not a blanket requirement to store all content in full. Which information is needed, who may see it and how long it is retained must be defined for the specific use.

Reviewing a language model's output must not be equated with validating a rule-based knowledge model. There, too, passing validation remains bound to the subject and limits of the review. A formally appropriate output format does not yet answer whether the maintenance instruction has been reproduced correctly in substance.

If a person is to act on this basis, they must be able to check relevant information and reject an unsuitable answer; only then is human oversight effective. For the choice of operating arrangement, the specific review question is therefore: which document passages does each service process, with which model version—and against what can the returned answer subsequently be checked?

Primary sources and further reading

[1] NIST: The NIST Definition of Cloud Computing. SP 800-145, September 2011, section 2, pp. 2–3. Original source

[2] Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020; version used: arXiv:2005.11401v4, 12 April 2021, Fig. 1 and section 2, pp. 2–3. Original source

[3] Karpukhin et al.: Dense Passage Retrieval for Open-Domain Question Answering. EMNLP 2020, pp. 6769–6781; here section 3.1, pp. 6770–6771. Original source

[4] BSI: Generative KI-Modelle: Chancen und Risiken für Industrie und Behörden. Version 2.0, 17 January 2025; pp. 36–37 and 48. Original source

[5] NIST: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. AI 600-1, July 2024; section 2.2, GV-6.1-007, GV-6.2-004, GV-6.2-007 and MS-2.5-003. Original source

[6] ISO/IEC 22989:2022, Information technology — Artificial intelligence — Artificial intelligence concepts and terminology, Edition 1, 2022; terms 3.1.17, 3.3.5, 3.3.7–3.3.8, 3.3.14–3.3.16. Paid standard. Original source.