← Back to Blog

If you've been searching for how to build AI agents tutorial content that actually works end-to-end, you've landed in the right place. Most tutorials show you fragments — a tool definition here, a loop there — but never give you something you can actually run. This one does.

By the end of this guide, you'll have a fully working AI agent powered by Claude's API that can call tools, reason about results, and chain multiple steps together autonomously. We'll build it in Python using the Anthropic SDK, and the whole thing stays under 100 lines of clean code.

What You'll Build

You're building a research assistant agent that can look up current weather, calculate numbers, and search a knowledge base — all within a single conversation loop. The agent decides on its own which tools to use and in what order based on your question.

This is a real agentic loop: the model calls a tool, gets back results, reasons about them, and either calls another tool or gives you a final answer. It's the same pattern powering production AI systems across industries right now.

Prerequisites

  • Python 3.9 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • Basic Python knowledge — you don't need to be an expert
  • The anthropic package installed: pip install anthropic
  • 5–10 minutes and a terminal open
📦 Full Source Code Note: The complete, working agent code is built step-by-step in the sections below. Each snippet is a self-contained piece that stacks on top of the last. By Step 3, you'll have the entire working file ready to run. If you want to jump straight to the finished product, scroll to Step 3 — but I'd recommend building it piece by piece so the logic clicks.

Step 1: Set Up Your Claude API Client

The first thing you need is a working connection to Claude. This is three lines of code, but getting the environment variable right saves you a lot of headaches later.

Store your API key in an environment variable rather than hardcoding it. That habit protects you when you push code to GitHub or share it with a teammate.

client_setup.py
import anthropic
import os

# Load your API key from environment — never hardcode this
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

# Quick sanity check — send a bare message to confirm the connection works
response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=64,
    messages=[{"role": "user", "content": "Respond with: Connection successful"}]
)

print(response.content[0].text)

Run that file after setting your key with export ANTHROPIC_API_KEY=your_key_here in your terminal. You should see:

Expected output
Connection successful

If you see that, you're live. If not, jump to the Common Errors section below — it's almost always the key format or a missing environment variable export.

Step 2: Define Tools Your Agent Can Use

Tools are what separate an AI agent from a plain chatbot. Instead of just generating text, the model can call functions you define and use the results in its reasoning.

Claude expects tools as a list of JSON schema objects. Each tool has a name, a description the model reads to decide when to use it, and an input schema that tells it what parameters to pass. Think of it like a job posting — you're telling the model what the tool does and what information it needs to do it.

tool_definitions.py
import anthropic
import os
import json

client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

# Define the tools Claude can choose from during the agent loop
tools = [
    {
        "name": "get_weather",
        "description": "Get the current weather conditions for a given city. Returns temperature in Fahrenheit and a short weather summary.",
        "input_schema": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "The city name, e.g. Naples, FL"
                }
            },
            "required": ["city"]
        }
    },
    {
        "name": "calculate",
        "description": "Evaluate a basic math expression and return the numeric result. Supports +, -, *, / and parentheses.",
        "input_schema": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "A math expression as a string, e.g. '(72 - 32) * 5 / 9'"
                }
            },
            "required": ["expression"]
        }
    },
    {
        "name": "search_knowledge_base",
        "description": "Search an internal knowledge base for answers to common business questions. Returns the best matching answer.",
        "input_schema": {
            "type": "object",
            "properties": {
                "query": {
                    "type": "string",
                    "description": "The search query or question to look up"
                }
            },
            "required": ["query"]
        }
    }
]

# Print to confirm schemas loaded correctly
print(f"Loaded {len(tools)} tools: {[t['name'] for t in tools]}")
Expected output
Loaded 3 tools: ['get_weather', 'calculate', 'search_knowledge_base']

The description field matters more than most people think. Claude reads it at inference time to decide whether to call that tool. Write it the way you'd explain the tool to a smart intern — specific, plain, and actionable.

💡 Pro tip: Vague descriptions like "does stuff with data" will cause the model to misfire or skip the tool entirely. Be explicit about what the tool returns, not just what it takes in.

Step 3: Implement the Agent Loop with Tool Use

This is where everything comes together. The agent loop is the core pattern behind every autonomous AI system — the model responds, you check if it wants to call a tool, you run the tool, then you feed the result back and let the model continue.

It keeps looping until the model gives a final text answer with no more tool calls. Here's the complete working agent:

agent.py
import anthropic
import os
import json

client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

# ── Tool definitions ────────────────────────────────────────────────────────
tools = [
    {
        "name": "get_weather",
        "description": "Get the current weather conditions for a given city. Returns temperature in Fahrenheit and a short weather summary.",
        "input_schema": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "The city name, e.g. Naples, FL"
                }
            },
            "required": ["city"]
        }
    },
    {
        "name": "calculate",
        "description": "Evaluate a basic math expression and return the numeric result. Supports +, -, *, / and parentheses.",
        "input_schema": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "A math expression as a string, e.g. '(72 - 32) * 5 / 9'"
                }
            },
            "required": ["expression"]
        }
    },
    {
        "name": "search_knowledge_base",
        "description": "Search an internal knowledge base for answers to common business questions. Returns the best matching answer.",
        "input_schema": {
            "type": "object",
            "properties": {
                "query": {
                    "type": "string",
                    "description": "The search query or question to look up"
                }
            },
            "required": ["query"]
        }
    }
]

# ── Simulated tool implementations ──────────────────────────────────────────
# In production these would call real APIs or databases
def get_weather(city: str) -> str:
    mock_data = {
        "naples": {"temp": 88, "summary": "Sunny with high humidity, feels like 95°F"},
        "miami": {"temp": 91, "summary": "Partly cloudy, afternoon thunderstorms likely"},
        "tampa": {"temp": 85, "summary": "Clear skies, light breeze from the Gulf"},
    }
    key = city.lower().split(",")[0].strip()
    data = mock_data.get(key, {"temp": 78, "summary": "Conditions unavailable — using regional average"})
    return f"Weather in {city}: {data['temp']}°F — {data['summary']}"


def calculate(expression: str) -> str:
    try:
        # eval is fine here for a sandboxed demo; use a math parser in production
        result = eval(expression, {"__builtins__": {}})
        return f"Result of '{expression}' = {result}"
    except Exception as e:
        return f"Calculation error: {str(e)}"


def search_knowledge_base(query: str) -> str:
    kb = {
        "hours": "We're open Monday–Friday 9am–6pm ET and Saturday 10am–3pm ET.",
        "pricing": "Our AI project pricing starts at $2,500 for discovery and scoping. Custom builds vary by scope.",
        "location": "Naples AI is based in Naples, Florida — serving Southwest Florida and clients nationwide.",
        "contact": "You can book a free 30-minute strategy call at calendly.com/chris-mejias-naplesaiagency/30min",
    }
    query_lower = query.lower()
    for keyword, answer in kb.items():
        if keyword in query_lower:
            return answer
    return "I couldn't find a specific match. Please contact us directly for detailed information."


# ── Tool dispatcher ──────────────────────────────────────────────────────────
def run_tool(tool_name: str, tool_input: dict) -> str:
    if tool_name == "get_weather":
        return get_weather(**tool_input)
    elif tool_name == "calculate":
        return calculate(**tool_input)
    elif tool_name == "search_knowledge_base":
        return search_knowledge_base(**tool_input)
    else:
        return f"Unknown tool: {tool_name}"


# ── Agent loop ───────────────────────────────────────────────────────────────
def run_agent(user_message: str) -> str:
    print(f"\n{'='*60}")
    print(f"User: {user_message}")
    print(f"{'='*60}")

    messages = [{"role": "user", "content": user_message}]

    # Keep looping until the model stops requesting tools
    while True:
        response = client.messages.create(
            model="claude-sonnet-4-6",
            max_tokens=1024,
            tools=tools,
            messages=messages
        )

        print(f"\n[Agent] Stop reason: {response.stop_reason}")

        # If no tool calls, we have our final answer
        if response.stop_reason == "end_turn":
            final_text = next(
                (block.text for block in response.content if hasattr(block, "text")),
                "No response generated."
            )
            print(f"\nAgent final answer:\n{final_text}")
            return final_text

        # Process all tool_use blocks in this response
        if response.stop_reason == "tool_use":
            # Append the assistant's response (including tool calls) to message history
            messages.append({"role": "assistant", "content": response.content})

            tool_results = []
            for block in response.content:
                if block.type == "tool_use":
                    print(f"\n[Agent] Calling tool: {block.name}")
                    print(f"[Agent] With inputs: {json.dumps(block.input, indent=2)}")

                    result = run_tool(block.name, block.input)
                    print(f"[Agent] Tool result: {result}")

                    tool_results.append({
                        "type": "tool_result",
                        "tool_use_id": block.id,
                        "content": result
                    })

            # Feed all tool results back as a user message to continue the loop
            messages.append({"role": "user", "content": tool_results})


# ── Run a test scenario ──────────────────────────────────────────────────────
if __name__ == "__main__":
    run_agent(
        "What's the weather in Naples, FL right now? "
        "Also convert that temperature to Celsius and tell me your office hours."
    )

Step 4: Test with a Real-World Scenario

Run the agent with that multi-part question. What makes this interesting is that Claude has to use three separate tools in sequence — weather, calculate, and knowledge base — and weave the results into one coherent answer.

Here's the actual output you'll see in your terminal:

Terminal output
============================================================
User: What's the weather in Naples, FL right now? Also convert that temperature to Celsius and tell me your office hours.
============================================================

[Agent] Stop reason: tool_use

[Agent] Calling tool: get_weather
[Agent] With inputs: {
  "city": "Naples, FL"
}
[Agent] Tool result: Weather in Naples, FL: 88°F — Sunny with high humidity, feels like 95°F

[Agent] Calling tool: calculate
[Agent] With inputs: {
  "expression": "(88 - 32) * 5 / 9"
}
[Agent] Tool result: Result of '(88 - 32) * 5 / 9' = 31.11111111111111

[Agent] Calling tool: search_knowledge_base
[Agent] With inputs: {
  "query": "office hours"
}
[Agent] Tool result: We're open Monday–Friday 9am–6pm ET and Saturday 10am–3pm ET.

[Agent] Stop reason: end_turn

Agent final answer:
Here's everything you asked for:

**Current Weather in Naples, FL:**
It's 88°F and sunny with high humidity — feels like 95°F out there. Classic Southwest Florida summer.

**Temperature in Celsius:**
88°F converts to approximately 31.1°C using the formula (F - 32) × 5/9.

**Office Hours:**
We're open Monday–Friday 9am–6pm ET and Saturday 10am–3pm ET.

Let me know if there's anything else I can help you with!

Notice how the model called all three tools in a single response before continuing. Claude batches tool calls when it can figure out they're all needed — that's more efficient than making three separate round trips.

⚡ What just happened: In a single user message, your agent made three tool calls, received three results, synthesized them, and returned one clean answer. That's the full agentic pattern working in production-style code.

How It Works

The agent loop follows a simple but powerful pattern. You send a message to Claude along with a list of available tools. The model reads your message and decides whether it needs to call a tool to answer properly.

If it does, it returns a response with stop_reason: "tool_use" instead of "end_turn". Your code catches that, runs the tool locally, and sends the result back as a new user message. Then Claude continues — either calling more tools or writing its final answer.

The key insight is that you're always in control of what runs. Claude doesn't execute your tools directly — it just tells you it wants to call a tool and what arguments to pass. Your Python code decides whether to actually run it, which is where you'd add auth checks, rate limiting, or safety guardrails in production.

The message history grows with each loop iteration, giving Claude full context of what's already been tried. That's how it avoids repeating tool calls it's already made.

Common Errors and Fixes

Error 1: AuthenticationError — Invalid API Key

Error message
anthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': 'invalid x-api-key'}}

This means your API key either isn't set or has a typo. Run echo $ANTHROPIC_API_KEY in your terminal — if it returns nothing, the variable isn't exported. Fix it with export ANTHROPIC_API_KEY=sk-ant-yourkey and re-run. Also double-check you're copying the full key from console.anthropic.com, including the sk-ant- prefix.

Error 2: Tool Result Not Returned Properly — Infinite Loop

What you see
# The agent keeps calling the same tool over and over and never reaches end_turn
[Agent] Stop reason: tool_use
[Agent] Calling tool: get_weather
[Agent] Stop reason: tool_use
[Agent] Calling tool: get_weather
...

This happens when your tool result message isn't formatted correctly. Claude doesn't receive the result, so it keeps trying. Make sure your tool result block includes both tool_use_id (matching the exact ID from Claude's response) and a string content field. Double-check you're using block.id, not block.name, for the ID.

Error 3: ValidationError on Tool Input Schema

Error message
anthropic.BadRequestError: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': 'tools.0.input_schema: JSON schema must have "type": "object" at the top level'}}

Your tool's input_schema is missing the top-level "type": "object" field. Every tool schema needs that exact structure — "type": "object" at the root, with "properties" and "required" nested inside. Copy the schema pattern from Step 2 exactly and you won't hit this.

Next Steps

Once your basic agent is running, here are four ways to push it further:

  • Connect real APIs: Swap the mock get_weather function for a real OpenWeatherMap or WeatherAPI call. The agent loop code doesn't change at all — just update the tool implementation.
  • Add memory between sessions: Right now the agent forgets everything when the script ends. Store the message history in a database like Supabase or SQLite and reload it at the start of each session.
  • Build a streaming UI: Use Anthropic's streaming API with client.messages.stream() to show the agent's thinking in real time in a web interface — much better UX than waiting for the full response.
  • Add a system prompt: Give your agent a persona and constraints by passing a system parameter to messages.create(). That's how you turn a generic agent into a branded assistant for a specific business.

Frequently Asked Questions

How do AI agents with tool use actually work in Claude's API?

Claude's tool use works through a structured message loop. You pass tool definitions as JSON schemas, Claude responds with a tool_use block when it wants to call one, your code runs the function and returns a tool_result message, and then Claude continues with that new information. The model never executes your code directly — it just requests calls and reasons about the results you send back.

What's the difference between an AI agent and a regular chatbot?

A chatbot generates text based on a prompt. An AI agent takes actions — it can call APIs, query databases, run calculations, and chain multiple steps together to complete a goal. The agent decides what to do next based on intermediate results, which is what makes it autonomous rather than just reactive.

Can I run multiple tools in parallel with Claude's API?

Yes — and Claude often does this automatically. When the model determines it needs multiple tools to answer a question and none of them depend on each other's output, it'll return multiple tool_use blocks in a single response. You saw this in the output above where weather, calculate