Open-Source · Technologies
Ollama
Open models on your hardware, one command away.
Ollama runs open models on your own hardware with a single command, and serves them through an API your applications can use like any hosted model. It is the standard for local AI. We size the hardware, deploy the models and wire Ollama into real workloads, so sovereignty stops being a slide and becomes a server.
LOCAL MODELS · EXAMPLE SERVER
MODEL
USE
SERVED VIA
STATUS
general-chat
Everyday assistant
Local API
SERVING
coding-model
Developer tooling
Local API
SERVING
embedding-model
Search and RAG
Local API
SERVING
gpu-headroom
Capacity planning
Monitored
REVIEW
oversized-model
Beyond the hardware
Swapping
RISK
Sample server
serving · review · risk
In plain terms
Some data should never leave the building. Now the model does not have to either.
Ollama puts capable open models on hardware you control. Here is what changes.
Without it
- Sensitive data sent to cloud APIs because there was no alternative
- Per token bills that grow with every workload
- AI unavailable the moment the internet is not
- Vendor roadmaps deciding what your stack can do
With it
- Models running where the data already lives
- Steady workloads at hardware cost, not per token
- AI that works in air gapped and offline environments
- A stack you version, swap and control outright
What CG TECH can do with Ollama
The work, broken into the parts that matter.
The data stays home
Prompts, documents and outputs never leave your infrastructure, which is the clean answer for regulated data, client confidentiality and air gapped environments.
prompts that never leave the building
Capable models, no licence desk
The major open model families run through one interface: general assistants, coding models and embeddings, swapped with a command. We benchmark them on your actual tasks.
the right open model, proven on your tasks
Local models, standard interface
Ollama serves an OpenAI compatible endpoint, so existing applications and agent frameworks point at your server instead of a cloud API with minimal change.
your apps, pointed at your hardware
Sized to the job, not the hype
Model quality tracks memory and GPU, so we size hardware to the workload, quantise where it is sensible and tell you plainly when a task still belongs on a frontier cloud model.
hardware sized honestly
How an engagement runs
From cloud only to a local capability, step by step.
01
Discover
We map the workloads and data that suit local models.
02
Size
Hardware, models and quantisation matched to the tasks.
03
Build
Ollama deployed, models benchmarked, workloads wired in.
04
Handover
Monitoring, updates and a tuning rhythm your team owns.
Questions we hear a lot
Common questions about Ollama
What is Ollama?
Ollama runs open models on your own hardware with a single command, and serves them through an API your applications can use like any hosted model. It is the common standard for running local AI.
Are local models good enough for real work?
For many workloads, yes: drafting, summarising, classification, coding assistance and retrieval run well on current open models. Frontier cloud models still win the hardest reasoning, and we route honestly between both.
What hardware do we need?
It depends on the models and volume. Capable setups start with a single GPU workstation and scale to dedicated servers. We size it to your tasks during discovery rather than selling the biggest box.
What does it cost to run?
Hardware, power and upkeep instead of per token fees. At steady volume that trade often wins, and it is predictable. We model the comparison plainly before you commit.
How does it fit with our cloud AI?
As one layer of a hybrid. Sensitive and steady workloads run locally; the hardest problems go to frontier models. The same applications can use both, and we design that routing.
How is Ollama different from LM Studio?
LM Studio is a desktop app for one person to try models. Ollama is the piece you run on a server so that applications and teams can use models over an API. Most projects use both: LM Studio to evaluate, Ollama to deploy.
Can more than one person use the same server?
Yes, that is the point of serving models through an API. How many people it comfortably supports at once depends on the model size and the hardware, which we size before you buy.
How do we update a model?
You pull the new version and switch to it, keeping the old one until you are satisfied. Test the new model against a fixed set of your own questions, so you can see whether it actually got better for your work.
Ready when you are
Want AI on your own hardware? Let us talk.
A discovery session maps your workloads, your risks and your quick wins. You keep the plan either way.
What to expect
- A consultant replies within 4 business hours
- Session booked to understand your requirements
- We will provide you with a fixed price quote