Anthropic · Technologies
Claude API
Claude models, built into your products and workflows.
The Claude API puts the Claude model family behind your own software: frontier Opus for the hard problems, Sonnet for the everyday agentic work and Haiku for speed and cost. With 1 million tokens of context on current models, it is built for long documents and serious agents. We design, build and run Claude workloads properly.
MODEL WORKLOADS · EXAMPLE PLATFORM
WORKLOAD
MODEL
COST PROFILE
STATUS
Contract-Review
Opus
Premium
LIVE
Support-Agent
Sonnet
Balanced
LIVE
Doc-Extraction
Haiku
Low cost
LIVE
Nightly-Batch
Batch API
Discounted
REVIEW
legacy-model-calls
Retired path
Unknown
MIGRATE
Sample platform
live · review · migrate
In plain terms
The most expensive model is rarely the right one.
API value comes from matching models to jobs and engineering the costs. Here is what changes.
Without the work
- Every call routed to the biggest model by default
- Token bills that climb without explanation
- Long documents chopped up to fit small context windows
- Model retirements discovered when production breaks
With the work
- Each workload on the cheapest model that does the job
- Caching and batching cutting the bill before it arrives
- Whole documents and codebases handled in one call
- Model upgrades planned on Anthropic’s cadence, not against it
What CG TECH can do with the Claude API
The work, broken into the parts that matter.
Capability matched to cost
Opus 4.8 for frontier reasoning, Sonnet 5 for agentic work at a sharp price, Haiku 4.5 for volume. We benchmark your actual tasks and route each workload to the model that earns its price.
capability matched to cost
Whole documents, not fragments
Current Claude models carry a 1 million token context window at standard rates, so contracts, codebases and case files fit in one call instead of a stitching exercise.
the whole document in one call
A unit cost you can defend
Prompt caching, the Batch API for non urgent work, token budgets and routing rules. The difference between a casual integration and an engineered one is usually the bill.
a unit cost you can defend
Your cloud, your call
Direct through the Claude Platform, or through Amazon Bedrock, Google Cloud and Microsoft Foundry, which means Claude fits inside a Microsoft estate cleanly. We recommend per environment, honestly.
Claude in the cloud you already run
How an engagement runs
From a raw API key to an engineered workload, step by step.
01
Discover
We map your use cases, volumes and quality bar.
02
Design
Model routing, cost controls and evaluation designed up front.
03
Build
The workload shipped with monitoring and fallbacks in place.
04
Handover
Runbooks, dashboards and an upgrade rhythm your team owns.
Questions we hear a lot
Common questions about the Claude API
What is the Claude API?
The Claude API puts the Claude model family behind your own software: Opus for the hard problems, Sonnet for everyday agentic work and Haiku for speed and cost. Current models carry up to a million tokens of context, which is what makes long documents and serious agents practical.
Which model do we actually need?
Usually a mix. Everyday tasks run well on Haiku or Sonnet, and only the genuinely hard steps justify Opus. We benchmark on your real tasks rather than guessing.
What does the API cost?
It is usage based, priced per million tokens and varying by model. The honest answer depends on volumes and design, so we model it during discovery, including the savings from caching and batching.
Is our API data used to train models?
No. Anthropic does not train on API business data by default.
Can we run it inside our Microsoft or AWS environment?
Yes. Claude models are available through Microsoft Foundry, Amazon Bedrock and Google Cloud as well as directly from Anthropic, so billing and governance can sit where your cloud already lives.
What does a million tokens of context actually mean?
Roughly a very large set of documents held in mind at once, in the order of thousands of pages. It matters because it removes a lot of the awkward engineering that used to be needed to feed a model just the right excerpt, and it makes agents that work across a whole codebase or case file realistic.
How do we stop costs running away?
Send the cheapest model that does the job, cache the stable parts of prompts, and watch usage from the first week rather than the first invoice. Prompt caching alone routinely takes more than half off a production bill.
What happens when a new model is released?
They arrive often, which is a good thing if you are set up for it. Keep the model as a configuration setting rather than hard-coded, and hold a small test set so you can check a new one against your own work before switching.
Ready when you are
Building on Claude? Let us talk.
A discovery session maps your use cases, your risks and your quick wins. You keep the plan either way.
What to expect
- A consultant replies within 4 business hours
- Session booked to understand your requirements
- We will provide you with a fixed price quote