An AI API is a web endpoint that lets your code send text to a language model and receive the model's reply as data. You pay per token, meaning per small chunk of text processed, and you never touch the model's weights or the hardware it runs on. Every major model lab sells one, Anthropic and Google among them.
The examples use Anthropic's Claude API, whose documentation we checked every number against, and the shape carries over to other providers.
How an AI API works: one request, one reply
Your program sends an HTTP request containing a conversation, and the provider returns the model's next message. That is the whole contract.
The request is a JSON document that names the model and carries a list of messages, each marked as coming from the user or the assistant. A max_tokens field caps the length of the reply. An optional system prompt, meaning standing instructions for the whole conversation, sets the role and the rules. Your API key, a secret string that ties each request to your account, travels in a header.
With Anthropic's Python SDK, a complete call fits on one screen:
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
system="Answer in two sentences or fewer.",
messages=[{"role": "user", "content": "What does an AI API do?"}],
)
for block in response.content:
if block.type == "text":
print(block.text)
print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens)
The reply arrives as a list of content blocks, so read them by type: current models can return a thinking block ahead of the text. A stop_reason of end_turn means the model finished, and max_tokens means your cap cut it off. The usage field is your meter.
Anthropic's guide to the Messages API states that the API is stateless, so you send the full conversation history with every request. Each request arrives like a colleague back from a long holiday: well rested and fully briefed on nothing.
That has a price attached. Turn twenty of a chat resends turns one to nineteen, so a long conversation costs more with every exchange.
Choosing an AI API provider
You can reach the same model through more than one door, and the door decides who bills you and which data terms apply.
| Route | Examples | Fits when |
|---|---|---|
| The model maker's own API | Claude API, Gemini API | You want new models and features first |
| A cloud platform | Amazon Bedrock, Google Cloud, Microsoft Foundry | Your billing and access control already live there |
| A model you host yourself | Open-weight models on your own hardware | Data must stay on your network |
Self-hosting swaps the per-token bill for a hardware bill. Our guide to running a local LLM on your own machine covers how much memory that takes.
Benchmarks will not settle the choice. Every launch announcement claims first place on some leaderboard, and most of them are telling the truth, which says more about leaderboards than about models. Run twenty of your own prompts through two candidates and compare the answers.
AI API pricing: what a million tokens buys
AI APIs bill by the token, with separate rates for the text you send and the text you get back.
Anthropic's API pricing page puts one token at about four characters, or three quarters of a word in English. Prices are quoted per million tokens:
| Model | Input, per million tokens | Output, per million tokens | Anthropic suggests it for |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $5 | Simple tasks |
| Claude Sonnet 5 | $2 | $10 | Most production workloads |
| Claude Opus 5 | $5 | $25 | The most complex reasoning |
The same page works through a support desk example: 10,000 tickets at roughly 3,700 tokens each come to about $37 on Claude Haiku 4.5. Two details decide whether your own bill looks that tidy.
The first is thinking. Current models such as Claude Sonnet 5 reason before answering by default, and Anthropic's migration guide notes that thinking tokens are billed as output even when the text is never returned. The effort setting controls how much reasoning you pay for.
The second is the tokenizer, the component that splits text into tokens. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text, so a lower per-token price does not guarantee a cheaper request.
Two levers cut the bill without touching quality. Prompt caching bills repeated reads of a long prefix, such as a system prompt, at a tenth of the input price on most models. The Batch API runs requests that can wait at half price. If you are choosing between the API and a chat plan, our breakdown of Claude.ai pricing plans and their usage limits compares the two.
What is LLM temperature, and why is it disappearing?
Temperature is a setting that controls how much randomness goes into a model's word choices. Low values make answers focused and repeatable, and high values make them varied.
The model scores every candidate for the next token. At a low temperature it almost always takes the top scorer; at a higher one it samples more widely. Anthropic's Messages API reference gives a range of 0.0 to 1.0 and warns that even 0.0 is not fully deterministic.
The same reference marks temperature as deprecated: Claude models released after Opus 4.6 reject any value other than the default 1.0 with a 400 error, the HTTP code for a malformed request. Paste a 2024 snippet with temperature=0.2 into a new project and you finally get perfectly deterministic behaviour from an LLM, since it fails the same way on every call.
On those models the effort parameter sets how hard the model thinks, and the prompt carries the tone. Ask for alternatives when you want variety, and spell out the format when you want consistency.
Build LLM apps with tool use and structured output
Two API features turn single calls into an application: tool use and structured output.
Tool use, also called function calling, lets you describe functions your code can run, each with a name and a JSON schema for its inputs. When a function would help, the model replies with a tool_use block naming it and its arguments. Your code runs the function and sends the output back in a tool_result block, and the model answers with real data in hand.
The model never executes anything itself. It hands you a neatly completed request form and waits, which makes it the most polite colleague you will ever have and the least productive one if nobody reads the forms.
Structured output covers the other half. You supply a JSON schema and the API constrains the reply to match it, so your code always gets something it can parse.
Put the two in a loop and you have an agent:
- Send the conversation and your tool definitions to the model.
- If the reply requests tools, run them and append the results.
- Repeat until the reply ends with
end_turn.
Every agent framework wraps some version of that loop. Agentic AI Engineering covers designing the tools and system prompts that keep it on task.
Your first AI API project, step by step
Treat the API key like a password. Keep it in an environment variable or a secrets manager, and never ship it inside a web page or a mobile app, where anyone can extract it and spend your money. Browser code should call a small backend of your own that holds the key.
Plan for limits. Anthropic caps requests and tokens per minute by usage tier, and going over returns a 429 status code that the official SDKs retry with backoff. Our guide to the Claude rate exceeded error and its fixes explains when retrying cannot help.
Then work through the rest in order:
- Make one call with the official SDK and print
usageto see what a request costs. - Write the system prompt before the code around it, since it shapes the output more than any parameter.
- Start with the cheapest model that passes on twenty real examples, and move up only where it fails.
- Log token usage from day one.
That last step turns a surprise invoice into a trend you saw coming. When a project grows past a single prompt, AI Engineering follows the same path from prompt design to multi-agent systems and fine-tuned models.