In short
An internal AI assistant is a ChatGPT-style tool where employees ask questions about company documents in plain language and get answers with the source attached. It is for companies whose knowledge is scattered across folders and people’s heads. CybUP prepares the documents, builds the RAG pipeline and permission-aware access, chooses a cloud or on-premises model, and puts KVKK data leakage controls in place.
What is an internal AI assistant?
An internal AI assistant is a ChatGPT-style chat interface that answers employees’ questions using the company’s own documents as its source. The difference is that it draws on your procedures, product documentation and policies rather than general internet knowledge. Under each answer it shows which document, and which section of it, the answer came from, so the employee can click through and check.
Picture a new production planner at a 120-person manufacturer who asks: “Who signs off quality on subcontracted production?” Today the answer lives either in an old Word file on a shared drive or in a senior colleague’s head. The assistant finds the relevant procedure, lists the steps briefly and links to the document.
What is RAG, and how does the assistant answer from your documents?
RAG (retrieval-augmented generation) is the method in which the model first finds the relevant passages in your company documents and then builds its answer on them. Documents are split into small chunks, each chunk is turned into a vector that represents its meaning, and the vectors are stored in a vector database. When a question comes in, the most relevant chunks are retrieved and passed to the model with the instruction “answer only from these texts”.
With this approach there is no need to retrain the model on company data. When a procedure changes, only that document is reprocessed, and the assistant uses the new information from the next question onwards. The real effort is on the document side: making scanned PDFs machine-readable, weeding out old versions, deciding which of two conflicting documents on the same subject is current. Build an assistant without that clean-up and all it does is spread the mess faster.
How does the assistant respect user permissions?
A permission-aware assistant answers each user only from documents they can already see in the source system. The salary spreadsheet in the HR folder, contracts in accounting or board minutes must not become visible to everyone through the assistant. So when documents are processed, the list of who may access each chunk is stored with it, and every search is filtered by the user’s group memberships.
We take user identity from your existing directory service, usually by setting up single sign-on with Active Directory or Microsoft Entra ID. When an employee changes department or leaves, their access to the assistant changes along with the directory. A separate user list means forgotten accounts.
“Implement fine-grained access controls and permission-aware vector and embedding stores. Ensure strict logical and access partitioning of datasets in the vector database to prevent unauthorized access between different classes of users or different groups.”
Cloud or on-premises AI assistant?
Cloud models are more capable and quicker to set up; on-premises models keep data inside the building. The choice depends on how sensitive the documents are, whether they contain personal data and how many people will use the assistant. With an enterprise cloud service, data is processed under the provider’s contractual terms, so those terms, and the region where the data is processed, need reading.
A hybrid setup is also possible. General procedures and product documentation can be answered by a cloud model, while documents containing personal data or trade secrets go to a model running in-house. We cover the hardware and installation side of the in-house option on our local LLM deployment page. Either way, the vector database and document store can sit on your server or in your own cloud subscription.
How do you manage KVKK and data leakage risks with an AI assistant?
Staff pasting customer lists, contracts or personnel details into public AI chat tools on personal accounts is one of the most common uncontrolled data leaks in companies today. An internal assistant moves that need into a channel you can audit. The Personal Data Protection Authority’s guide to generative AI and personal data protection (in Turkish) looks at data processing risks across the lifecycle of these systems, and KVKK, Türkiye’s Personal Data Protection Law (Law No. 6698), requires data controllers to take the technical and administrative measures needed for an appropriate level of security.
On the technical side, the OWASP Top 10 for LLM Applications is a good control framework. Its entries on sensitive information disclosure, prompt injection and system prompt leakage apply directly to an internal assistant. Even a line such as “ignore previous instructions” hidden inside a document can try to change how the model behaves.
- Permission filtering applied in the retrieval layer, not left to the model
- Questions and answers logged together with who asked them
- Separate rules, or a separate model, for document types that contain personal data
- A defined retention period for the logs
- The model provider’s data use terms documented
“LLMs, especially when embedded in applications, risk exposing sensitive data, proprietary algorithms, or confidential details through their output.”
How does an internal AI assistant project run?
The project starts with one department and a limited set of documents. First we gather that department’s 30 to 50 most frequently asked questions, together with the documents that hold the answers. That question list doubles as the test set for checking whether the assistant is working correctly.
The assistant opens to a pilot group; we track answer accuracy, how often it cites a source and whether it says “I don’t know” when it should. If the results are good, other departments’ documents are added. To keep documents current, we set up regular synchronisation from the shared drive or document management system, usually as an n8n workflow. The first meeting and review are free.
What you receive
- A chat interface with single sign-on that cites its sources
- A cleaned and classified document set and a vector database
- Permission-aware search filtered by user group
- The chosen model (cloud, on-premises or hybrid) and its connection configuration
- Question and answer logs, retention rules and a data flow document
- A pilot evaluation report with test set results
How we work
- 1
Free review
We assess the target department, document sources, number of users and data sensitivity.
- 2
Document preparation
Documents are collected, old versions are weeded out and permissions are mapped.
- 3
Infrastructure and model
The vector database, retrieval layer and chosen model are set up and connected to your directory service.
- 4
Pilot and measurement
The assistant opens to a pilot group, and accuracy is measured against the test set and real questions.
- 5
Rollout and maintenance
Other departments are added; synchronisation, updates and log reviews continue.
Frequently asked questions
How is this different from staff using ChatGPT?
General chat tools don’t know your documents, and when staff use them on personal accounts you can’t control where the data goes. An internal assistant answers from your documents, according to each user’s permissions, and cites its sources. The usage logs stay inside the company too.
Will our data be used to train the model?
With RAG we don’t train the model on your data; documents are only supplied as context at the moment of answering. If a cloud model is used, we check and document whether the provider’s enterprise terms allow data to be used for training. With an on-premises model, the data never leaves your server.
Which file types can it read?
Word, PDF, Excel, PowerPoint, text files and wiki pages are the usual sources. Scanned PDFs need OCR first, and the quality depends on the scan. Documents that are mostly tables need separate handling.
What if the assistant gives a wrong answer?
The assistant shows its source with every answer and is set up to say so plainly when the documents don’t cover a question. Even so, language models can make mistakes, which is why we tell users from the start to check the source before any critical decision. During the pilot we collect wrong answers and fix them.
How many users can it support?
What matters is less the number of users than how many ask questions at the same time, and which model is chosen. With a cloud model, scale depends on the provider; with an on-premises model, on the hardware. We design around expected usage.
How much does an internal AI assistant cost?
We send a written quote once we have reviewed the scope; the review is free. Cloud model usage charges, or the hardware for an on-premises option, are shown as separate items.
Does it work with SharePoint or our file server?
Yes. Documents can be synchronised from these sources on a schedule, and folder permissions can be carried over to the assistant. The connection method depends on the type of source. We check your current setup during the review.