In short
Local LLM deployment means running open-weight large language models (LLMs) on a company’s own GPU server, using software such as Ollama or vLLM. It is for companies that don’t want to, or can’t, send their data to the cloud. CybUP sizes the model and hardware, sources a GPU server if needed, and handles installation, RAG integration, security and updates.
What is a local LLM, and when does it make sense?
A local LLM is a language model whose weights you download and run on your own hardware; prompts and responses never go to an outside service. This has become possible because more and more model developers now publish their weights (open-weight models). Each model has its own licence and the terms for commercial use vary, so reading the licence is part of how we choose a model.
A local model makes sense in three situations: the data involves personal data or trade secrets and you don’t want it sent to a service abroad; usage is heavy and steady enough that ongoing cloud charges justify investing in hardware; or you need to work on an isolated network with limited internet access. For occasional work that needs the most capable model available, the cloud is usually more practical.
To be frank, models small enough to run in-house won’t match the largest cloud models on every task. For well-defined jobs such as summarising documents, classification, field extraction and answering questions from documents, they are often good enough. We show you this before installation by testing on a sample of your own data.
Ollama vs vLLM: what is the difference?
Ollama makes installing and managing models easy and suits serving a small number of users from a single server; vLLM is an inference server built to handle many concurrent requests with high throughput. Ollama’s Linux documentation covers installation, running it as a systemd service and support for NVIDIA and AMD GPUs. Because Ollama offers OpenAI-compatible API endpoints, many existing applications can switch to a local model just by changing the endpoint URL.
According to the vLLM documentation, vLLM is a library that manages memory efficiently with PagedAttention, batches incoming requests continuously and includes an OpenAI-compatible API server. For an internal assistant used by dozens of people at once, vLLM is the more efficient option. We suggest Ollama for small installations and trials and vLLM for heavy use; because we write applications against the OpenAI-compatible API, switching later is straightforward.
“Connect OpenAI clients to Ollama. Ollama supports a subset of the OpenAI API.”
What GPU server do you need for a local LLM?
GPU sizing comes down mainly to the model’s parameter count and the precision it runs at (quantisation). Roughly speaking, all of the model weights need to fit in GPU memory (VRAM). At 16-bit precision each parameter takes 2 bytes; 8-bit and 4-bit quantisation cut that to a half and a quarter. On top of the weights you need extra memory for concurrent requests and long contexts.
An example: suppose a 50-person law firm wants a mid-sized model for questions over its documents, with a few people asking at once. A single-GPU server can handle that. If the same model has to summarise call centre notes all day for a 300-person company, you need several GPUs or cards with more memory. We give you a firm size by running your sample workload in a test environment and measuring it.
The rest of the server matters too: enough system memory, fast NVMe storage for the model files, and the power supply and cooling the cards demand. We can source new enterprise GPU servers from abroad and install them; if you already have a suitable server, we assess it.
How do you connect company data to a local model?
The usual way to connect company documents to a local model is RAG: documents are split into chunks and converted into vectors with an embedding model, and when a question comes in the relevant chunks are retrieved and passed to the model. The embedding model runs locally as well, so no stage of document processing leaves your network. We describe the interface built on top of this on our internal AI assistant page.
A local model can also work inside automations, with no chat screen involved. An n8n workflow that classifies incoming email or extracts fields from PDFs, for example, can call the local model instead of a cloud one. For work involving personal data, that may well be the simplest solution.
How do you secure a local LLM server?
The model server should sit on your network as an internal service that only the applications using it can reach. According to the Ollama FAQ, Ollama listens only on 127.0.0.1, port 11434, by default; exposing it to the network means changing the bind address. When we change that setting, we put an authenticating reverse proxy in front of it and restrict access at the firewall to the servers that need it. A model API left open to the internet without authentication means strangers using your compute.
At the application layer we keep the risks in the OWASP Top 10 for LLM Applications in view: prompt injection, model output passed unchecked to other systems, and unbounded resource consumption. That is why per-request length limits, per-user rate limits and logging are part of every installation.
“Ollama binds 127.0.0.1 port 11434 by default. Change the bind address with the OLLAMA_HOST environment variable.”
How is a local LLM maintained and updated?
A local LLM is not something you install once and forget. GPU drivers, CUDA or ROCm components, Ollama or vLLM versions and the model itself all change over time. A driver update can stop the model server from starting, so we schedule changes and take a backup first.
When a new model version comes out, we don’t switch straight away. We run the test questions prepared during the pilot against the new model and compare the answers; if it does better, we switch. GPU temperature, memory use and response times are monitored and can feed into your existing monitoring system. On the operating system side, the usual Ubuntu Server maintenance routine applies.
What you receive
- A model and hardware sizing report based on your sample data
- Sourcing and installation of a new GPU server, if needed
- A model server running Ollama or vLLM with an OpenAI-compatible API
- RAG infrastructure with a local embedding model and a vector database
- Security configuration with authentication, network restrictions and logging
- A maintenance document covering updates, the test set and rollback steps
How we work
- 1
Free review
We assess your use cases, data sensitivity, number of users and current hardware.
- 2
Trial and sizing
Candidate models are tested on your sample data, and the GPU and server configuration is worked out.
- 3
Hardware and installation
The server is prepared or sourced; drivers, the model server and the models are installed.
- 4
Integration and security
RAG, applications and automations are connected; access restrictions and logging are switched on.
- 5
Maintenance
Driver, software and model updates are tested before they are applied, and performance is monitored.
Frequently asked questions
Is a local LLM as good as ChatGPT?
Not at everything. Models that can run on an in-house server often do well on defined tasks such as summarising documents, classification and answering questions from documents. We test this on your own data and show you the results before installation.
Will it run on our existing server?
Small models can run on a server without a GPU, but they are slow and not suitable for multiple users. If the chassis can take a suitable GPU and the power supply and cooling are up to it, an existing server can be considered. We check your hardware during the review.
Can you supply the GPU server?
Yes, if you wish. We source new enterprise GPU servers from abroad and install them. If you would rather buy from your own supplier, we prepare the technical specification.
Does a local LLM work without an internet connection?
Yes. Once the model and the required software have been downloaded, no internet connection is needed to run it. Updates are brought in separately, in a controlled way.
Can we use open-weight models commercially?
It depends on the model, as each has its own licence. Some are permissively licensed; others come with usage conditions. We read the licence with you and choose one that fits how you plan to use it.
How much does a local LLM deployment cost?
We send a written quote once we have reviewed the scope; the review is free. Hardware, installation and maintenance appear as separate items in the quote.
Is a local model enough for KVKK compliance?
No, not on its own. A local model stops data leaving the country, but KVKK, Türkiye’s Personal Data Protection Law (Law No. 6698), still requires you to deal with access rights, logs, retention periods and privacy notices. We document the technical side as part of the installation.