โ† Back to Articles
Messages API Tool Use Loop Streaming Prompt Caching Model Family Managed Agents Cost Reality Takeaways
AI Agents ยท LLM Architecture ยท API Design
๐Ÿค–
Anthropic / Claude

Anthropic Claude Architecture:
Messages API, Tool Use Loops,
and Managed Agents

Claude's API is a structured request/response surface: messages, tools, cache breakpoints, and a stop_reason you have to loop on. The managed Agents stack sits on top of that same loop.

๐Ÿ“… August 14, 2026 โœ Barnabas Waweru โฑ 14 min read ๐Ÿท AI Agents
ansi ยท wordmark ยท anthropic claude
 โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—
โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ•—  โ–ˆโ–ˆโ•‘โ•šโ•โ•โ–ˆโ–ˆโ•”โ•โ•โ•โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•
โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘     
โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ•โ• โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘     
โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘ โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—
โ•šโ•โ•  โ•šโ•โ•โ•šโ•โ•  โ•šโ•โ•โ•โ•   โ•šโ•โ•   โ•šโ•โ•  โ•šโ•โ•โ•šโ•โ•  โ•šโ•โ• โ•šโ•โ•โ•โ•โ•โ• โ•šโ•โ•     โ•šโ•โ• โ•šโ•โ•โ•โ•โ•โ•

 โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•—      โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ•—   โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•— โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—
โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•”โ•โ•โ•โ•โ•
โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—  
โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ•‘     โ–ˆโ–ˆโ•”โ•โ•โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘   โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ–ˆโ–ˆโ•”โ•โ•โ•  
โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—โ–ˆโ–ˆโ•‘  โ–ˆโ–ˆโ•‘โ•šโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•”โ•โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ•—
 โ•šโ•โ•โ•โ•โ•โ•โ•šโ•โ•โ•โ•โ•โ•โ•โ•šโ•โ•  โ•šโ•โ• โ•šโ•โ•โ•โ•โ•โ• โ•šโ•โ•โ•โ•โ•โ• โ•šโ•โ•โ•โ•โ•โ•โ•

The Messages API, A Reasoning Engine, Not a Chat Endpoint

The Messages API is a stateless POST that carries the whole conversation. The response is a typed content[] array, not a string. max_tokens is required. system is a top-level field, not a message role.

// Claude API, single request lifecycle
POST https://api.anthropic.com/v1/messages Headers: x-api-key: ANTHROPIC_API_KEY anthropic-version: 2023-06-01 content-type: application/json Body: model: "claude-opus-5" โ† required max_tokens: 4096 โ† required (no safe default) system: "..." โ† top-level, NOT a message role messages: [ {role, content} ] โ† full history on every call tools: [ {...} ] โ† optional tool definitions stream: true โ† optional SSE mode cache_control: { type: ... } โ† optional automatic caching Response: id: msg_01... type: message role: assistant content: [ TextBlock | ToolUseBlock | ThinkingBlock ] stop_reason: "end_turn" | "tool_use" | "max_tokens" usage: { input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens }
โšก Critical Gotcha

max_tokens is required. Omitting it throws a validation error, there is no safe default. system is a top-level parameter, not a message with role: "system". Passing it inside messages will error. Both of these burn junior developers on first contact.

๐Ÿ“จ

Messages API

POST /v1/messages, Stateless, full-history request. GA. The core of every Claude integration.

๐Ÿ“ฆ

Message Batches API

POST /v1/messages/batches, Async bulk processing at 50% cost reduction. GA. For eval pipelines, bulk classification, and offline processing.

๐Ÿ”ข

Token Counting API

POST /v1/messages/count_tokens, Count tokens before sending. GA. Essential for cache planning and rate-limit budgeting.

๐Ÿ“

Files API

POST /v1/files, Upload files once, reference across many calls. Beta. Eliminates re-encoding large documents on every request.

Content Block Array, the Real Response Unit

Claude's response is not a string. It is a content[] array of typed blocks. Accessing response.content[0].text directly will explode when the first block is a tool_use block, which happens whenever Claude decides to call a tool. Always check content[0].type before accessing fields.

  • TextBlock, { type: "text", text: string }. The most common block for plain responses.
  • ToolUseBlock, { type: "tool_use", id, name, input }. Indicates Claude wants to call a function. Your code must execute it and return results.
  • ThinkingBlock, { type: "thinking", thinking: string }. Extended reasoning trace, only on models and requests that enable it. Not visible in standard responses.

Tool Use Architecture, The Execution Loop

Tool use is commonly described as "function calling" but that framing misses the architecture. It is a multi-turn protocol where Claude signals intent, your application executes, and the result flows back into the conversation. The model itself never runs any code. You do. Claude just decides when and how to call.

// Tool taxonomy, client vs server
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ TOOL TYPES โ”‚ โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค โ”‚ CLIENT TOOLS (your code executes) โ”‚ โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ โ”‚ โ”‚ User-defined tools โ”‚ โ”‚ โ”‚ โ”‚ โ†’ pass input_schema, you handle call โ”‚ โ”‚ โ”‚ โ”‚ โ”‚ โ”‚ โ”‚ โ”‚ Anthropic-defined client schemas โ”‚ โ”‚ โ”‚ โ”‚ โ†’ bash, text_editor (you exec locally) โ”‚ โ”‚ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค โ”‚ SERVER TOOLS (Anthropic executes for you) โ”‚ โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ โ”‚ โ”‚ web_search_20260209 โ”‚ โ”‚ โ”‚ โ”‚ โ†’ Anthropic infra runs the search โ”‚ โ”‚ โ”‚ โ”‚ โ†’ cited results in same response โ”‚ โ”‚ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
1

Define tools with JSON Schema

Pass a tools[] array with name, description, and input_schema. The description is read by Claude, a vague description produces poor tool selection. Be explicit about when the tool should and should not be called.

2

Claude returns stop_reason: "tool_use"

The response content will contain a ToolUseBlock with { id, name, input }. The id is critical, it links the tool call to the tool result in the next turn.

3

You execute the tool

Run the function, hit the API, query the database, whatever the tool requires. This happens entirely in your application. Claude is waiting with its context intact.

4

Return tool_result back to Claude

Append the assistant message (Claude's full content array) and a user message with { type: "tool_result", tool_use_id, content }. The tool_use_id must match the id from step 2.

5

Loop until stop_reason: "end_turn"

Claude may call multiple tools in sequence or parallel. The loop continues until stop_reason is "end_turn". A single messages.create call is never enough for agentic tasks, build the while loop, not a one-shot call.

Tool Use Loop, Minimal TypeScript

const messages: MessageParam[] = [ { role: "user", content: "What's the status of order #12345?" } ]; let response = await client.messages.create({ model: "claude-sonnet-4-6", max_tokens: 1024, tools, messages, }); // The loop, not optional for agentic tasks while (response.stop_reason === "tool_use") { const toolBlock = response.content.find(b => b.type === "tool_use"); if (!toolBlock || toolBlock.type !== "tool_use") break; const result = await runTool(toolBlock.name, toolBlock.input); messages.push({ role: "assistant", content: response.content }); messages.push({ role: "user", content: [{ type: "tool_result", tool_use_id: toolBlock.id, // must match content: result, }], }); response = await client.messages.create({ model, max_tokens, tools, messages }); }
๐Ÿ”‘ Architectural Principle

The tool use loop is your responsibility, not Anthropic's. Claude is stateless between calls. The growing messages[] array is the entire agent state. Every tool call round-trip costs input tokens proportional to history length, this is where agentic costs compound. Design tools that return concise, structured results.

Streaming Architecture, SSE and Content Block Deltas

Long outputs without streaming mean one thing: gateway timeouts and users staring at a blank screen. Claude's streaming mode uses Server-Sent Events (SSE) with a structured event sequence. It is not just chunked text, it is a typed event stream that includes tool use blocks, thinking deltas, and usage stats.

// SSE event sequence for a streamed response
event: message_start โ†’ { type, message: {id, model, usage} } event: content_block_start โ†’ { index, content_block: {type, text} } event: content_block_delta โ†’ { index, delta: {type, text} } โ† text chunks event: content_block_delta โ†’ { index, delta: {type, text} } โ† ... event: content_block_stop โ†’ { index } event: message_delta โ†’ { delta: {stop_reason}, usage } event: message_stop โ†’ {} // For tool_use blocks: event: content_block_start โ†’ { content_block: {type: "tool_use", id, name} } event: content_block_delta โ†’ { delta: {type: "input_json_delta", partial_json} } event: content_block_stop โ†’ {}

When Streaming is Not Optional

  • Code generation, responses easily exceed 2000 tokens, and Vercel/Netlify function timeouts at 10โ€“30s will clip non-streamed responses.
  • Documentation synthesis, same issue; streaming gives users readable progress immediately.
  • Multi-turn agents, stream each turn so users know the agent is working, not frozen.
  • Extended thinking, thinking tokens can be very long. You almost certainly want to stream them or skip rendering them.
๐Ÿ

Python SDK

client.messages.stream() context manager. stream.text_stream for text-only iteration. Both sync and async supported.

๐ŸŸฆ

TypeScript SDK

client.messages.stream() with .on("text", cb) event handler. Typed event objects. Promise-based final message via stream.finalMessage().

๐Ÿ”Œ

Raw SSE

Any HTTP client that handles text/event-stream. Use when you need the raw event sequence for custom routing (e.g., proxy to WebSocket).

๐Ÿ› 

Streaming + Tool Use

Works together. Tool use blocks arrive as streamed JSON deltas. Accumulate input_json_delta events, parse on content_block_stop.

Prompt Caching, The Most Underused Cost Lever

Prompt caching is the single biggest cost optimization available to heavy Claude users, and most teams are not using it. The idea is simple: mark a stable prefix of your prompt, system context, tool definitions, RAG documents, and Anthropic will cache it server-side. Cache reads cost approximately one-tenth of normal input token pricing.

Two Caching Modes

Automatic caching, Add a single top-level cache_control: { type: "ephemeral" } to the request. Anthropic applies the breakpoint automatically and moves it forward as the conversation grows. Best for multi-turn conversations where the accumulated history should be cached between turns.

Explicit breakpoints, Place cache_control directly on individual content blocks for precise control. Supports up to four breakpoints per request. Use when you have a known stable prefix (large system prompt, product catalog, codebase context) followed by dynamic per-user content.

// Prompt structure, cache break placement
system: [ { type: "text", text: LARGE_SYSTEM_PROMPT, // ~4000 tokens, stable cache_control: { type: "ephemeral" } // โ† breakpoint 1 }, { type: "text", text: PRODUCT_CATALOG, // ~10000 tokens, stable cache_control: { type: "ephemeral" } // โ† breakpoint 2 (max 4) } ] messages: [ { role: "user", content: USER_QUESTION } // โ† volatile, never cached ] // Response usage tells you if caching hit: usage.cache_creation_input_tokens // > 0 on first call (write) usage.cache_read_input_tokens // > 0 on subsequent calls (hit) // Both 0 = prefix too short, cache miss, or dynamic content before breakpoint
โš  Cache Invalidation Rules

The cache prefix must be byte-for-byte identical across calls. Any dynamic value before the cache_control breakpoint, a timestamp, a user ID, unsorted JSON keys, will invalidate the cache on every call. Never interpolate volatile data before the last breakpoint. Also: minimum cacheable prefix is ~1024โ€“4096 tokens depending on model. Shorter prefixes silently skip caching with no error, just cache_creation_input_tokens: 0.

TTL Options

  • { type: "ephemeral" }, Default 5-minute TTL. Enough for bursty request windows where many users share the same system prompt.
  • { type: "ephemeral", ttl: "1h" }, Extended TTL for longer-lived contexts. Useful when your system prompt is stable over a session that spans 30โ€“60 minutes.

The Model Family, Pick the Right Blade

The 2026 Claude model family has four tiers in active production, plus invitation-only models under Project Glasswing. The right model is not the most capable one, it is the most capable one that fits your latency, cost, and quality requirements for the specific task.

Model ID Sweet Spot Tradeoff Status
Claude Fable 5 claude-fable-5 Long-running agents, hardest reasoning, enterprise work Highest cost, highest latency, not for latency-sensitive paths GA
Claude Opus 5 claude-opus-5 Complex agentic coding, multi-step code review, deep reasoning High cost; use where quality justifies it GA
Claude Sonnet 4.6 claude-sonnet-4-6 Everyday code gen, PR reviews, support agents, the default workhorse Best cost/quality balance for most production workloads GA
Claude Haiku 4.5 claude-haiku-4-5-20251001 High-volume classification, streaming autocomplete, cheap routing Lower reasoning depth; use only for well-defined, simple tasks GA
Claude Mythos 5 claude-mythos-5 Research frontier tasks, Project Glasswing invitation-only Not generally available; requires approved access INVITE
๐ŸŽฏ Routing Strategy

A practical two-tier setup: Sonnet for the hot path (user-facing, latency-sensitive, most requests), Opus or Fable for the cold path (background agents, complex code review, deep analysis). Use the Token Counting API before the hot path to route requests that will need heavy reasoning to the cold path automatically.

โ˜

AWS Bedrock

All GA models available. IAM-integrated auth. Same API surface, Bedrock billing. Best for teams already on AWS with VPC data residency requirements.

๐ŸŸฆ

Google Cloud Vertex AI

GA models on Vertex. Google Cloud IAM auth. Useful when the rest of the stack is GCP and you want unified billing and compliance.

๐ŸชŸ

Microsoft Foundry

GA models via Azure-integrated Microsoft Foundry. Entra ID auth. For teams in the Microsoft enterprise ecosystem.

๐Ÿ”‘

Direct API

Latest models and features first. Anthropic billing and support. x-api-key header or Workload Identity Federation for keyless auth.

Managed Agents Stack, The Emerging Control Plane

Building a tool use loop yourself is the low-level API. Anthropic now ships a higher-level managed agent runtime that handles state, sandboxed execution environments, and versioned agent definitions. It is still Beta as of August 2026. If it sticks, they own the execution layer, not just the model.

// Managed Agents stack, four layered APIs
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ YOUR APPLICATION โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ–ผ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ SESSIONS API /v1/sessions โ”‚ โ”‚ Stateful agent sessions in managed cloud sandboxes โ”‚ โ”‚ Event stream: GET /v1/sessions/{id}/events/stream โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ–ผ โ–ผ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ AGENTS API โ”‚ โ”‚ ENVIRONMENTS API โ”‚ โ”‚ /v1/agents โ”‚ โ”‚ /v1/environments โ”‚ โ”‚ Versioned agent โ”‚ โ”‚ Sandbox templates, โ”‚ โ”‚ definitions with โ”‚ โ”‚ OS, tools, deps โ”‚ โ”‚ skills + tools โ”‚ โ”‚ available in sessions โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ–ผ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ SKILLS API /v1/skills โ”‚ โ”‚ Custom capabilities attached โ”‚ โ”‚ to agent definitions โ”‚ โ”‚ (versioned, reusable) โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Agents API, Versioned Agent Definitions

An Agent definition bundles a model, a system prompt, a set of tools and skills, and environment configuration into a versioned, reusable object. Once defined, you run sessions against it rather than reconstructing the full context on every call.

  • Versioned, Agent definitions are immutable snapshots. Promote from staging to production like a container image.
  • Composable, Attach Skills (custom capabilities) to an Agent. Skills are the unit of reusable behavior in the Anthropic managed stack.
  • Decoupled from execution, The Agent definition says what the agent can do. The Session says when it runs and for whom.

Sessions API, Stateful Execution in Managed Sandboxes

A Session is a running instance of an Agent in a managed cloud sandbox. Anthropic handles state persistence, tool routing, and environment setup. The client streams events via GET /v1/sessions/{id}/events/stream, an SSE endpoint that emits agent actions, tool calls, and output in real time.

  • Stateful, The session maintains conversation history server-side. No more threading a growing messages[] array through your application.
  • Sandboxed, Code execution, file access, and external calls happen inside Anthropic's managed infrastructure, not your Lambda function.
  • Observable, The event stream gives you full visibility into agent actions, tool calls, and intermediate reasoning. Build audit logs from it.
๐Ÿงญ Claude Code Architecture

Claude Code (Anthropic's CLI) is itself built on this stack. When you run Claude Code as an agent, it uses the same tool use loop, session state model, and skills system described here, just packaged into a CLI workflow with bash, text_editor, and file system client tools. The Claude Code Max subscription uses subscription billing rather than API token pricing, which radically changes the economics for high-volume interactive workloads.

API Endpoint Purpose GA / Beta
Files POST /v1/files Upload once, reference across many calls. Eliminates repeated encoding. BETA
Skills POST /v1/skills Create versioned, reusable agent capabilities. BETA
Agents POST /v1/agents Define reusable, versioned agent configurations. BETA
Sessions POST /v1/sessions Run stateful agent sessions in managed cloud sandboxes. BETA
Environments POST /v1/environments Configure sandbox templates for agent sessions. BETA

Cost Reality Check

// Cost Reality Check

Claude API billing is per input and output token. Costs compound fast in agentic loops because every tool call round-trip re-sends the full conversation history as input tokens. A 10-tool-call agent session with a 4000-token system prompt can easily consume 50,000+ input tokens per run.

  • Prompt caching is your biggest lever. If your system prompt is 4000 tokens and you get 1000 requests/day, caching alone can cut daily input costs by 70โ€“80% after the first cache write per TTL window.
  • Batch API for non-real-time work. 50% cost reduction with async processing. If your eval pipeline or classification job is not user-facing, it should be using batches.
  • Model routing saves money. Haiku is dramatically cheaper per token than Sonnet. Routing simple classification, intent detection, and tagging to Haiku while keeping Sonnet for generation is a 5โ€“10x cost reduction on the routed volume.
  • Token counting before generation. Use POST /v1/messages/count_tokens to gate expensive requests. If a request will consume >10k tokens, decide whether to proceed, warn the user, or summarize the context first.
  • Claude Code Max subscription changes the math. For teams using Claude Code interactively or in CI pipelines, the subscription model has no per-token billing. High-volume interactive agentic workloads often pay less per effective token on Max than on API pricing.
  • Keep tool results concise. Every tool result gets fed back as input tokens on the next call. A tool that returns 5000 tokens of raw JSON when a 50-token summary would do is burning your token budget silently.

// Key Takeaways

1

The Messages API is stateless by design. Your application owns the conversation state. Build the growing messages[] array correctly and never assume Claude remembers previous calls.

2

Tool use is a loop, not a single call. Claude signals intent via stop_reason: "tool_use". You execute. You return tool_result with the matching tool_use_id. Repeat until "end_turn". One call is never enough for agentic tasks.

3

Prompt caching is the single most impactful cost optimization. A stable system prompt longer than ~1024 tokens should always have a cache_control breakpoint. Verify it's working by checking usage.cache_read_input_tokens.

4

Stream everything that could be long. Vercel and Netlify function timeouts will clip long non-streamed responses. Use client.messages.stream() for code gen, docs, and any agentic output that might exceed 60 seconds.

5

Model routing is not optional at scale. Use Haiku for classification and routing, Sonnet for everyday generation, Opus for complex agents, and Fable for the hardest long-running work. Never default to the largest model for tasks that don't need it.

6

The Managed Agents stack is Anthropic's play to own the agent runtime, not just the model. Agents, Sessions, Skills, and Environments are Beta today, but track them. If they stabilize, the case for building your own agent state management weakens significantly.

7

Tool result size is a hidden cost driver. Every tool result comes back as input tokens. Design tools to return the minimum information Claude needs, not raw API dumps. A well-formatted 100-token summary beats a 5000-token JSON blob every time.

Sources: Anthropic API Overview ยท Tool Use Guide ยท Prompt Caching ยท Models Overview ยท Streaming Guide