CG TECH

Open-Source · Technologies

Ollama

Open models on your hardware, one command away.

Ollama runs open models on your own hardware with a single command, and serves them through an API your applications can use like any hosted model. It is the standard for local AI. We size the hardware, deploy the models and wire Ollama into real workloads, so sovereignty stops being a slide and becomes a server.

LOCAL MODELS · EXAMPLE SERVER

MODEL

USE

SERVED VIA

STATUS

general-chat

Everyday assistant

Local API

SERVING

coding-model

Developer tooling

Local API

SERVING

embedding-model

Search and RAG

Local API

SERVING

gpu-headroom

Capacity planning

Monitored

REVIEW

oversized-model

Beyond the hardware

Swapping

RISK

Sample server

serving · review · risk

In plain terms

Some data should never leave the building. Now the model does not have to either.

Ollama puts capable open models on hardware you control. Here is what changes.

Without it

  • Sensitive data sent to cloud APIs because there was no alternative
  • Per token bills that grow with every workload
  • AI unavailable the moment the internet is not
  • Vendor roadmaps deciding what your stack can do

With it

  • Models running where the data already lives
  • Steady workloads at hardware cost, not per token
  • AI that works in air gapped and offline environments
  • A stack you version, swap and control outright

What CG TECH can do with Ollama

The work, broken into the parts that matter.

How an engagement runs

From cloud only to a local capability, step by step.

01

Discover

We map the workloads and data that suit local models.

02

Size

Hardware, models and quantisation matched to the tasks.

03

Build

Ollama deployed, models benchmarked, workloads wired in.

04

Handover

Monitoring, updates and a tuning rhythm your team owns.

Questions we hear a lot

Common questions about Ollama

What is Ollama?

Ollama runs open models on your own hardware with a single command, and serves them through an API your applications can use like any hosted model. It is the common standard for running local AI.

Are local models good enough for real work?

For many workloads, yes: drafting, summarising, classification, coding assistance and retrieval run well on current open models. Frontier cloud models still win the hardest reasoning, and we route honestly between both.

What hardware do we need?

It depends on the models and volume. Capable setups start with a single GPU workstation and scale to dedicated servers. We size it to your tasks during discovery rather than selling the biggest box.

What does it cost to run?

Hardware, power and upkeep instead of per token fees. At steady volume that trade often wins, and it is predictable. We model the comparison plainly before you commit.

How does it fit with our cloud AI?

As one layer of a hybrid. Sensitive and steady workloads run locally; the hardest problems go to frontier models. The same applications can use both, and we design that routing.

How is Ollama different from LM Studio?

LM Studio is a desktop app for one person to try models. Ollama is the piece you run on a server so that applications and teams can use models over an API. Most projects use both: LM Studio to evaluate, Ollama to deploy.

Can more than one person use the same server?

Yes, that is the point of serving models through an API. How many people it comfortably supports at once depends on the model size and the hardware, which we size before you buy.

How do we update a model?

You pull the new version and switch to it, keeping the old one until you are satisfied. Test the new model against a fixed set of your own questions, so you can see whether it actually got better for your work.

Ready when you are

Want AI on your own hardware? Let us talk.

A discovery session maps your workloads, your risks and your quick wins. You keep the plan either way.

What to expect

Scroll to Top