← Back to Blog

If you've searched for how to build AI agents tutorial, you've probably hit the same wall I did — most guides are either too abstract or show you toy examples that fall apart the moment you try something real. This guide fixes that. You'll build a working autonomous agent using Python and the Anthropic SDK that can reason through problems, call tools, and loop until it finishes a task — not just generate a single response.

What You'll Build

By the end of this tutorial, you'll have a production-ready AI agent that uses Claude Sonnet 4.6 to autonomously answer questions by calling real tools — a web search simulator and a calculator. The agent runs a loop, decides which tools to use, executes them, and keeps going until it has a final answer.

This is the same architecture we use at Naples AI when building custom agent systems for local businesses. It's minimal, but it's real — and you can extend it to production in an afternoon.

📦 Full Source Code
The complete working code is built step-by-step in the sections below. Each step adds one piece of the agent — by Step 4 you'll have the full system running. No placeholder logic, no pseudocode. Copy each block in order and you're done.

Prerequisites

  • Python 3.10 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • Basic familiarity with Python classes and functions
  • The anthropic Python SDK installed (pip install anthropic)
  • A terminal and a code editor — nothing else needed

Step 1: Set Up Claude API and Anthropic SDK

First, install the Anthropic SDK and confirm your API key is working. I always test the connection with a bare-minimum call before building anything on top of it — saves a lot of debugging later.

Set your API key as an environment variable so it never touches your source code. On Mac/Linux run export ANTHROPIC_API_KEY=your_key_here. On Windows use set ANTHROPIC_API_KEY=your_key_here.

setup_check.py
import anthropic
import os

# Confirm the SDK can reach the API before building anything on top of it
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=64,
    messages=[{"role": "user", "content": "Say hello in one sentence."}]
)

print(response.content[0].text)
        

If your key is set correctly you'll see something like: "Hello! It's great to meet you." If you get an AuthenticationError, your key isn't loaded — double-check the environment variable name and restart your terminal.

💡 Tip
Never hardcode your API key directly in source files. If you push to GitHub with a key in the code, Anthropic will auto-revoke it and you'll need a new one.

Step 2: Define Your Agent's Tools and Capabilities

Tools are how Claude takes action in the world. You define them as a JSON schema — Claude reads that schema and decides when and how to call each tool during a conversation. Think of it like a function signature with a description Claude can actually understand.

For this tutorial, I'm defining two tools: a search tool and a calculator. In a real production system these would hit actual APIs, but the agent logic is identical regardless of what the tool does under the hood.

tools.py
import anthropic
import json
import math

# Tool definitions tell Claude what capabilities it has and how to invoke them
TOOLS = [
    {
        "name": "web_search",
        "description": (
            "Search the web for current information about a topic. "
            "Use this when you need facts, news, or data you don't already know."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "query": {
                    "type": "string",
                    "description": "The search query string"
                }
            },
            "required": ["query"]
        }
    },
    {
        "name": "calculator",
        "description": (
            "Evaluate a mathematical expression and return the result. "
            "Use this for any arithmetic, algebra, or numeric computation."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "A valid Python math expression, e.g. '(15 * 4) / 2 + math.sqrt(9)'"
                }
            },
            "required": ["expression"]
        }
    }
]


def execute_tool(tool_name: str, tool_input: dict) -> str:
    """Run the requested tool and return the result as a string."""

    if tool_name == "web_search":
        query = tool_input["query"]
        # Simulated search results — swap this for a real search API in production
        simulated_results = {
            "naples florida population": "Naples, Florida has a population of approximately 22,000 in the city proper, with the greater Naples metro area exceeding 390,000 residents as of 2025.",
            "anthropic claude models": "Anthropic's current model lineup includes Claude Sonnet 4-6, Claude Opus 4, and Claude Haiku 3-5, released throughout 2025 and 2026.",
            "average restaurant profit margin": "The average restaurant profit margin is 3–9% for full-service restaurants and 6–9% for fast-casual concepts.",
        }
        # Return closest match or a default message
        for key, value in simulated_results.items():
            if any(word in query.lower() for word in key.split()):
                return value
        return f"Search results for '{query}': No specific data found. Please refine your query."

    elif tool_name == "calculator":
        expression = tool_input["expression"]
        try:
            # Allow math module functions inside expressions
            result = eval(expression, {"__builtins__": {}}, {"math": math})
            return str(result)
        except Exception as e:
            return f"Calculation error: {e}"

    return f"Unknown tool: {tool_name}"
        

The input_schema field is what Claude actually reads to understand what arguments to pass. If your schema is vague, Claude will guess — and it'll sometimes guess wrong. Be specific in your descriptions.

Step 3: Implement the Main Agent Class with Tool Use

Now we build the agent class itself. This is the core of the whole system — it holds the conversation history, sends messages to Claude, and handles the tool call responses that come back. The class design keeps everything in one place so the orchestration logic in Step 4 stays clean.

agent.py
import anthropic
import json
import os
from tools import TOOLS, execute_tool


class ClaudeAgent:
    """
    A single-agent system that uses Claude Sonnet 4-6 with tool use.
    Maintains conversation history across turns in the run loop.
    """

    def __init__(self, system_prompt: str = None):
        self.client = anthropic.Anthropic(
            api_key=os.environ.get("ANTHROPIC_API_KEY")
        )
        self.model = "claude-sonnet-4-6"
        self.tools = TOOLS
        self.conversation_history = []

        # Default system prompt instructs the agent to use tools proactively
        self.system_prompt = system_prompt or (
            "You are a helpful research assistant with access to web search and a calculator. "
            "When answering questions, use your tools to get accurate, up-to-date information. "
            "Think step-by-step and always verify numbers with the calculator tool. "
            "Be concise in your final answer."
        )

    def _send_message(self, user_message: str) -> anthropic.types.Message:
        """Append the user message to history and call the Claude API."""

        self.conversation_history.append({
            "role": "user",
            "content": user_message
        })

        response = self.client.messages.create(
            model=self.model,
            max_tokens=4096,
            system=self.system_prompt,
            tools=self.tools,
            messages=self.conversation_history
        )

        return response

    def _handle_tool_calls(self, response: anthropic.types.Message) -> None:
        """
        Process all tool_use blocks in the response.
        Appends the assistant's tool calls and our tool results back into history
        so Claude sees the full context on the next turn.
        """

        # Record what the assistant decided to do
        self.conversation_history.append({
            "role": "assistant",
            "content": response.content
        })

        # Build a tool_result block for every tool_use block in the response
        tool_results = []
        for block in response.content:
            if block.type == "tool_use":
                print(f"  → Calling tool: {block.name}({json.dumps(block.input)})")
                result = execute_tool(block.name, block.input)
                print(f"  ← Result: {result}\n")

                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": result
                })

        # Feed all tool results back as a single user turn
        self.conversation_history.append({
            "role": "user",
            "content": tool_results
        })
        

Notice that _handle_tool_calls appends both the assistant's tool-use blocks and the tool results back into conversation_history. This is the part most beginners miss. Claude needs to see its own tool calls and the results in the history to continue reasoning correctly.

⚠️ Common Mistake
If you forget to append the assistant's tool_use content blocks to history before sending tool results, Claude throws a 400 Bad Request error. The history must be: user → assistant (with tool_use) → user (with tool_result). That exact order matters.

Step 4: Create the Run Loop and Orchestration Logic

This is where the agent becomes autonomous. The run loop keeps calling Claude, handling tool calls, and feeding results back in — until Claude returns a final text response instead of another tool call. That's how you know it's done.

agent.py (continued — add this method to ClaudeAgent)
    def run(self, user_message: str, max_iterations: int = 10) -> str:
        """
        Main agent loop. Runs until Claude produces a final text answer
        or until max_iterations is reached to prevent infinite loops.
        """

        print(f"\n{'='*60}")
        print(f"USER: {user_message}")
        print(f"{'='*60}\n")

        response = self._send_message(user_message)
        iterations = 0

        # Keep looping as long as Claude wants to use tools
        while response.stop_reason == "tool_use" and iterations < max_iterations:
            iterations += 1
            print(f"[Iteration {iterations}] Claude is using tools...\n")
            self._handle_tool_calls(response)

            # Send an empty continuation — Claude reads the tool results from history
            response = self.client.messages.create(
                model=self.model,
                max_tokens=4096,
                system=self.system_prompt,
                tools=self.tools,
                messages=self.conversation_history
            )

        # Extract the final plain-text answer from the response
        final_answer = ""
        for block in response.content:
            if hasattr(block, "text"):
                final_answer += block.text

        # Add the final assistant response to history for multi-turn conversations
        self.conversation_history.append({
            "role": "assistant",
            "content": response.content
        })

        print(f"\nAGENT: {final_answer}")
        return final_answer
        

The max_iterations guard is not optional — it's essential. Without it, a misconfigured tool can cause the agent to loop forever and burn through your API credits. Ten iterations is generous for most tasks; you can lower it for simpler use cases.

Step 5: Test with Real-World Scenarios

Now let's wire everything together and run the agent against a few real questions. This file is your entry point — run it and you'll see the full reasoning loop in your terminal.

main.py
import anthropic
import json
import os
import math
from tools import TOOLS, execute_tool
from agent import ClaudeAgent


def main():
    agent = ClaudeAgent()

    # Test 1: Research question that needs web search
    agent.run(
        "What is the population of Naples, Florida, "
        "and how many restaurants would serve 10% of that population daily "
        "if each restaurant seats 80 people and turns tables 3 times a day?"
    )

    # Test 2: Pure calculation to verify the calculator tool works
    agent.run("What is the square root of 1764 multiplied by 15?")

    # Test 3: Multi-tool question requiring both search and math
    agent.run(
        "What is the average restaurant profit margin, "
        "and if a Naples restaurant does $1.2 million in annual revenue, "
        "what is the expected annual profit at the midpoint margin?"
    )


if __name__ == "__main__":
    main()
        

When you run python main.py, you'll see output that looks like this:

sample_output.txt
============================================================
USER: What is the population of Naples, Florida, and how many restaurants would
serve 10% of that population daily if each restaurant seats 80 people and turns
tables 3 times a day?
============================================================

[Iteration 1] Claude is using tools...

  → Calling tool: web_search({"query": "naples florida population"})
  ← Result: Naples, Florida has a population of approximately 22,000 in the city
     proper, with the greater Naples metro area exceeding 390,000 residents as of 2025.

[Iteration 2] Claude is using tools...

  → Calling tool: calculator({"expression": "390000 * 0.10 / (80 * 3)"})
  ← Result: 162.5

AGENT: Naples, Florida has a metro population of about 390,000 people.
To serve 10% of that population (39,000 people) daily, where each restaurant
seats 80 people and turns tables 3 times a day (240 covers per day),
you would need approximately 163 restaurants.

============================================================
USER: What is the square root of 1764 multiplied by 15?
============================================================

[Iteration 1] Claude is using tools...

  → Calling tool: calculator({"expression": "math.sqrt(1764) * 15"})
  ← Result: 630.0

AGENT: The square root of 1764 is 42, and 42 multiplied by 15 equals 630.
        

How It Works: Agent Decision Flow Explained

Here's what's actually happening under the hood on every loop iteration. Claude doesn't "run" your tools — it just tells you which tool to run and with what arguments. Your Python code runs the tool, hands the result back, and Claude decides what to do next.

The decision flow looks like this:

  1. User sends a message → appended to conversation_history
  2. Claude responds → either with a tool_use block (stop_reason: "tool_use") or a final text answer (stop_reason: "end_turn")
  3. If tool_use: your code runs the tool, appends both the assistant's tool call and your tool result to history, then calls Claude again
  4. If end_turn: extract the text, print it, done

The conversation history is the agent's working memory. Every tool call, every result, every response — it's all in there. Claude reads the full history on every API call, which is how it maintains context across multiple tool uses in a single task.

🧠 Key Insight
Claude doesn't have persistent memory between separate ClaudeAgent instances. If you want the agent to remember things across sessions, you need to save and reload self.conversation_history to a database or file. That's a common upgrade we add for production systems.

Common Errors and Fixes

Error 1: anthropic.BadRequestError — messages must alternate between user and assistant roles

This is the most common mistake when building the tool loop. It happens when you send tool results without first recording the assistant's tool_use response in history.

# WRONG — sending tool results without the assistant's prior response in history
conversation_history.append({"role": "user", "content": tool_results})

# RIGHT — always append the assistant response first, THEN the tool results
conversation_history.append({"role": "assistant", "content": response.content})
conversation_history.append({"role": "user", "content": tool_results})
        

Error 2: anthropic.AuthenticationError — invalid x-api-key

Your API key isn't loading. The environment variable name is case-sensitive and the terminal session matters — setting it in one terminal tab doesn't carry to another.

# Check if the key is actually available before running the agent
import os

api_key = os.environ.get("ANTHROPIC_API_KEY")
if not api_key:
    raise ValueError(
        "ANTHROPIC_API_KEY environment variable not set. "
        "Run: export ANTHROPIC_API_KEY=your_key_here"
    )

# Then pass it explicitly to the client
client = anthropic.Anthropic(api_key=api_key)
        

Error 3: Agent loops forever / hits max_iterations and stops mid-task

This usually means your tool is returning an error string instead of a useful result — so Claude keeps trying. Log every tool call result during development so you can see what's going wrong.

# Add this inside execute_tool() to debug tool failures during development
def execute_tool(tool_name: str, tool_input: dict) -> str:
    print(f"[DEBUG] Tool: {tool_name} | Input: {tool_input}")  # log every call
    result = ""

    if tool_name == "calculator":
        try:
            result = str(eval(tool_input["expression"], {"__builtins__": {}}, {"math": math}))
        except Exception as e:
            result = f"Error: {e}"  # return the error as a string, not an exception

    print(f"[DEBUG] Result: {result}")  # log every result
    return result
        

Next Steps: Scaling to Multi-Agent Systems

A single agent with a handful of tools is a solid starting point. Here's where to take it next once you have the basics working:

1. Add Real Tool Integrations

Swap the simulated search for a real API — Brave Search, Serper, or Tavily all have Python SDKs and free tiers. Add a tool that reads from a Google Sheet or queries a database. That's when the agent becomes genuinely useful for business workflows.

2. Build a Multi-Agent System with Orchestration

Anthropic's multi-agent architecture lets one "orchestrator" agent spin up specialized "subagent" instances — a research agent, a writing agent, a data agent. Each handles its own domain, and the orchestrator coordinates the output. We use this pattern for complex automation workflows at Naples AI.

3. Add Persistent Memory

Store conversation_history in a database like PostgreSQL or Redis between sessions. This lets your agent remember past interactions, user preferences, and task context — which is the difference between a demo and a production product.

4. Wrap It in an API with FastAPI

Put your ClaudeAgent class behind a FastAPI endpoint so any frontend or other service can talk to it over HTTP. Add a task queue like Celery if you need the agent to run long jobs in the background without blocking the request. That's the standard production setup.

Frequently Asked Questions

How do I add more tools to my Claude API agent in Python?