OpenAI · Technologies
Whisper
Every recording, turned into text you can use.
Whisper is OpenAI’s speech recognition model, open source and self hostable, with newer hosted transcription models available through the API. Meetings, calls and field recordings become searchable, summarisable text. We build transcription pipelines and, where sovereignty matters, we run them on your own hardware.
TRANSCRIPTION JOBS · EXAMPLE WEEK
JOB
SOURCE
PIPELINE
STATUS
Meeting-Minutes
Meeting exports
Automated
DONE
Call-Recordings
Support line
Automated
DONE
Field-Notes
Mobile audio
Automated
DONE
Podcast-Captions
Marketing
In review
REVIEW
personal-uploads
Unknown app
Ungoverned
RISK
Sample week
done · review · risk
In plain terms
Hours of audio. Minutes of reading.
Recordings are where information goes to hide. Transcription sets it free. Here is what changes.
Without it
- Meeting recordings nobody ever replays
- Call insights locked inside audio files
- Staff retyping interviews and site notes by hand
- Audio uploaded to random free tools with your data attached
With it
- Every recording transcribed automatically
- Calls and meetings searchable like documents
- Summaries and actions generated from the transcript
- A governed pipeline, on infrastructure you choose
What CG TECH can do with Whisper
The work, broken into the parts that matter.
From backlog to pipeline
Batch processing turns archives and daily recordings into text automatically, with speaker aware formatting where the source allows it.
the audio backlog, cleared
The right deployment for your data
Whisper is open source, so it can run entirely on your own hardware when recordings must not leave your network. The hosted API offers newer models and zero infrastructure. We advise honestly and run both.
the right deployment for your data
Text is the start, not the end
Transcripts feed summaries, minutes, action lists and searchable archives, chained with language models so the recording becomes an asset rather than a file.
recordings that answer questions
Trust, measured
Strong multilingual accuracy out of the box, tuned further with domain vocabulary, and human review inserted exactly where the stakes require it.
accuracy you have measured, not assumed
How an engagement runs
From audio backlog to searchable text, step by step.
01
Discover
We map your audio sources, volumes and sensitivity.
02
Design
Deployment choice, pipeline and review points designed to fit.
03
Build
The pipeline live, accuracy measured on your real audio.
04
Handover
Monitoring, costs and a tuning rhythm your team owns.
Questions we hear a lot
Common questions about Whisper
What is OpenAI Whisper?
Whisper is OpenAI’s speech recognition model. It is open source and can be self-hosted, and there are newer hosted transcription models available through the API. It turns meetings, calls and field recordings into text you can search and summarise.
How accurate is it?
Very good on clear speech and respectable in messy conditions, and it improves with domain vocabulary. We measure accuracy on your actual recordings during the pilot rather than quoting a brochure number.
Does it handle languages other than English?
Yes, dozens, including translation to English. We test the languages that matter to you before anything goes live.
What about privacy and sensitive recordings?
That is the deciding factor in deployment. Self hosted Whisper keeps audio entirely on your infrastructure, and the hosted API does not train on your business data. We match the deployment to the sensitivity.
What does transcription cost?
Self hosted costs are hardware and upkeep; hosted is priced per minute of audio. At volume the comparison gets interesting, and we model both honestly.
Why would we self-host it rather than use the API?
Data sovereignty, usually. If recordings contain health information, legal matters or anything that cannot leave your environment, running it on your own hardware answers that cleanly. Otherwise the hosted models are less work.
Does it handle Australian accents and place names?
Accents are generally handled well. Local place names, organisation names and industry terms are where errors show up, and that is fixable by giving it the vocabulary rather than accepting the default output.
What do businesses do with the transcripts?
Search them, summarise them, and pull actions out of them. The value is rarely the transcript itself. It is that six months of meetings and calls become something you can ask questions of.
Ready when you are
Sitting on hours of recordings? Let us talk.
A discovery session maps your audio, your risks and your quick wins. You keep the plan either way.
What to expect
- A consultant replies within 4 business hours
- Session booked to understand your requirements
- We will provide you with a fixed price quote