← Back to Blog

If you've searched for how to build AI agents tutorial and landed here, you're probably past the "what is an AI agent" stage and ready to actually build one. Good. This tutorial walks you through building a real autonomous agent using the Claude API and Python — one that can reason, use tools, and loop until it finishes a task on its own.

We built versions of this exact pattern at Naples AI for clients ranging from real estate agencies in Southwest Florida to manufacturing shops that needed automated data lookups and reporting. The pattern works. Let's get into it.

What You'll Build

You'll build a Python-based AI agent that uses Claude to answer multi-step questions by calling real tools — in this case, a weather lookup and a calculator. The agent decides which tools to call, calls them, reads the results, and keeps going until it has a complete answer.

By the end, you'll have a working agentic loop in about 50 lines of Python. You'll understand how tool use works with the Claude API, and you'll have a base you can extend into anything — a lead qualifier, a research assistant, or a process automation engine.

Prerequisites

  • Python 3.9 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • Basic Python knowledge — functions, classes, dictionaries
  • The anthropic SDK installed (pip install anthropic)
  • A terminal and a text editor or IDE
📦 Full Source Code
The complete, working code for this tutorial is broken into steps below so you can follow along and understand each piece. By the end of Step 5, you'll have the full agent assembled and ready to run. Copy each snippet in order and you'll have everything you need.

Step 1: Set Up Your Claude API Environment

First, install the Anthropic SDK if you haven't already. Open your terminal and run:

terminal
pip install anthropic

Next, set your API key as an environment variable. Don't hard-code it in your script — that's how keys get accidentally committed to GitHub.

terminal
export ANTHROPIC_API_KEY="your-api-key-here"

Now create a file called agent.py and start with the main agent class. This handles Claude client initialization and stores the conversation history the agent needs to reason across multiple steps.

agent.py
import anthropic
import json
import os

class ClaudeAgent:
    def __init__(self):
        # Initialize the Anthropic client using ANTHROPIC_API_KEY env variable
        self.client = anthropic.Anthropic()
        self.model = "claude-sonnet-4-6"
        self.messages = []

    def add_message(self, role: str, content):
        """Append a message to the running conversation history."""
        self.messages.append({"role": role, "content": content})

    def run(self, user_input: str) -> str:
        """
        Entry point: takes a user question, runs the agentic loop,
        and returns the final text answer.
        """
        self.add_message("user", user_input)
        return self._agentic_loop()

The messages list is what gives the agent memory within a session. Every user message, every tool result, and every Claude response gets appended here, so Claude always has the full context when deciding what to do next.

Step 2: Define Your Agent's Tools and Functions

Tools are how Claude interacts with the outside world. You define them as JSON schemas — Claude reads the schema to understand what the tool does and what parameters it needs. Then your Python code actually runs the tool when Claude asks for it.

Here we're defining two tools: a fake weather lookup and a simple calculator. In a real project, these would hit actual APIs or databases. The structure is identical regardless of what the tool does.

agent.py (add below the class definition)
# Tool definitions with JSON schema — Claude reads these to decide when to use each tool
TOOLS = [
    {
        "name": "get_weather",
        "description": (
            "Returns the current temperature and conditions for a given city. "
            "Use this when the user asks about weather or temperature in a location."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "The name of the city, e.g. 'Naples, FL'"
                }
            },
            "required": ["city"]
        }
    },
    {
        "name": "calculate",
        "description": (
            "Evaluates a simple mathematical expression and returns the result. "
            "Use this for any arithmetic the user needs."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "A math expression as a string, e.g. '(72 - 32) * 5 / 9'"
                }
            },
            "required": ["expression"]
        }
    }
]


def get_weather(city: str) -> str:
    """Simulated weather lookup. Replace with a real API call in production."""
    mock_data = {
        "naples": {"temp_f": 91, "conditions": "Sunny"},
        "miami": {"temp_f": 88, "conditions": "Partly Cloudy"},
        "new york": {"temp_f": 74, "conditions": "Overcast"},
    }
    key = city.lower().split(",")[0].strip()
    data = mock_data.get(key, {"temp_f": 75, "conditions": "Unknown"})
    return json.dumps({"city": city, "temperature_f": data["temp_f"], "conditions": data["conditions"]})


def calculate(expression: str) -> str:
    """Safely evaluates a math expression using only numeric operations."""
    try:
        # Restrict eval to math operations only — never eval raw user input in production
        allowed = {k: v for k, v in vars(__builtins__).items()
                   if k in ("abs", "round", "min", "max", "pow")}
        result = eval(expression, {"__builtins__": allowed})
        return json.dumps({"expression": expression, "result": result})
    except Exception as e:
        return json.dumps({"error": str(e)})


def execute_tool(tool_name: str, tool_input: dict) -> str:
    """Routes a tool call from Claude to the correct Python function."""
    if tool_name == "get_weather":
        return get_weather(**tool_input)
    elif tool_name == "calculate":
        return calculate(**tool_input)
    else:
        return json.dumps({"error": f"Unknown tool: {tool_name}"})

The description field in each tool definition is more important than it looks. Claude uses it to decide when to call the tool, so be specific. Vague descriptions lead to the agent calling the wrong tool or skipping one it should use.

Step 3: Implement the Agentic Loop with Tool Use

This is the core of a build AI agent Claude API Python setup — the agentic loop. Claude responds, you check if it wants to use a tool, you run the tool, you feed the result back, and you repeat until Claude stops asking for tools and gives a final answer.

Add the _agentic_loop method inside the ClaudeAgent class:

agent.py (inside ClaudeAgent class)
    def _agentic_loop(self) -> str:
        """
        Core agentic loop: sends messages to Claude, handles tool_use responses,
        feeds results back, and repeats until Claude returns a final text answer.
        """
        max_iterations = 10  # Safety cap to prevent infinite loops
        iteration = 0

        while iteration < max_iterations:
            iteration += 1

            response = self.client.messages.create(
                model=self.model,
                max_tokens=4096,
                tools=TOOLS,
                messages=self.messages
            )

            # Check why Claude stopped generating
            stop_reason = response.stop_reason

            if stop_reason == "end_turn":
                # Claude finished — extract and return the final text response
                final_text = next(
                    (block.text for block in response.content if hasattr(block, "text")),
                    "No response generated."
                )
                return final_text

            elif stop_reason == "tool_use":
                # Claude wants to call one or more tools
                # First, add Claude's full response (including tool_use blocks) to history
                self.add_message("assistant", response.content)

                # Collect all tool results before sending them back
                tool_results = []
                for block in response.content:
                    if block.type == "tool_use":
                        print(f"  → Agent calling tool: {block.name}({block.input})")
                        result = execute_tool(block.name, block.input)
                        tool_results.append({
                            "type": "tool_result",
                            "tool_use_id": block.id,
                            "content": result
                        })

                # Add all tool results in a single user message — Anthropic API requires this
                self.add_message("user", tool_results)

            else:
                # Unexpected stop reason — surface it rather than silently failing
                return f"Unexpected stop reason: {stop_reason}"

        return "Max iterations reached without a final answer."
⚠️ Important: When Claude calls multiple tools in one turn, you must return all tool results in a single user message. If you send them one at a time, the API will throw a validation error. The code above handles this correctly by collecting all results into tool_results before appending.

Step 4: Add Error Handling and Retry Logic

Real agents hit real problems — rate limits, network timeouts, malformed tool inputs. Without error handling, your agent crashes silently or returns garbage. Add this retry wrapper around your API call so transient failures don't kill a run.

agent.py (add imports at top, method inside class)
import anthropic
import json
import os
import time

class ClaudeAgent:
    def __init__(self):
        self.client = anthropic.Anthropic()
        self.model = "claude-sonnet-4-6"
        self.messages = []

    def add_message(self, role: str, content):
        self.messages.append({"role": role, "content": content})

    def _call_claude_with_retry(self, max_retries: int = 3) -> anthropic.types.Message:
        """
        Calls the Claude API with exponential backoff on rate limit or server errors.
        Raises the final exception if all retries are exhausted.
        """
        for attempt in range(max_retries):
            try:
                return self.client.messages.create(
                    model=self.model,
                    max_tokens=4096,
                    tools=TOOLS,
                    messages=self.messages
                )
            except anthropic.RateLimitError as e:
                if attempt == max_retries - 1:
                    raise
                wait = 2 ** attempt  # 1s, 2s, 4s
                print(f"  Rate limited. Retrying in {wait}s...")
                time.sleep(wait)
            except anthropic.APIStatusError as e:
                if attempt == max_retries - 1:
                    raise
                print(f"  API error {e.status_code}. Retrying...")
                time.sleep(2 ** attempt)

    def _agentic_loop(self) -> str:
        max_iterations = 10
        iteration = 0

        while iteration < max_iterations:
            iteration += 1

            response = self._call_claude_with_retry()
            stop_reason = response.stop_reason

            if stop_reason == "end_turn":
                final_text = next(
                    (block.text for block in response.content if hasattr(block, "text")),
                    "No response generated."
                )
                return final_text

            elif stop_reason == "tool_use":
                self.add_message("assistant", response.content)
                tool_results = []
                for block in response.content:
                    if block.type == "tool_use":
                        print(f"  → Agent calling tool: {block.name}({block.input})")
                        try:
                            result = execute_tool(block.name, block.input)
                        except Exception as e:
                            # Return the error as a tool result so Claude can handle it gracefully
                            result = json.dumps({"error": str(e)})
                        tool_results.append({
                            "type": "tool_result",
                            "tool_use_id": block.id,
                            "content": result
                        })
                self.add_message("user", tool_results)

            else:
                return f"Unexpected stop reason: {stop_reason}"

        return "Max iterations reached without a final answer."

    def run(self, user_input: str) -> str:
        self.add_message("user", user_input)
        return self._agentic_loop()

Wrapping tool execution in a try/except and returning errors as tool results is a pattern worth keeping. It lets Claude acknowledge the failure and try a different approach rather than crashing the whole run.

Step 5: Test Your Agent End-to-End

Now let's put it all together and run it. Add this at the bottom of agent.py so the script runs when you execute it directly.

agent.py (full final file)
import anthropic
import json
import os
import time


# ── Tool Definitions ────────────────────────────────────────────────────────

TOOLS = [
    {
        "name": "get_weather",
        "description": (
            "Returns the current temperature and conditions for a given city. "
            "Use this when the user asks about weather or temperature in a location."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "The name of the city, e.g. 'Naples, FL'"
                }
            },
            "required": ["city"]
        }
    },
    {
        "name": "calculate",
        "description": (
            "Evaluates a simple mathematical expression and returns the result. "
            "Use this for any arithmetic the user needs."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "A math expression as a string, e.g. '(72 - 32) * 5 / 9'"
                }
            },
            "required": ["expression"]
        }
    }
]


# ── Tool Implementations ─────────────────────────────────────────────────────

def get_weather(city: str) -> str:
    mock_data = {
        "naples": {"temp_f": 91, "conditions": "Sunny"},
        "miami": {"temp_f": 88, "conditions": "Partly Cloudy"},
        "new york": {"temp_f": 74, "conditions": "Overcast"},
    }
    key = city.lower().split(",")[0].strip()
    data = mock_data.get(key, {"temp_f": 75, "conditions": "Unknown"})
    return json.dumps({"city": city, "temperature_f": data["temp_f"], "conditions": data["conditions"]})


def calculate(expression: str) -> str:
    try:
        allowed = {k: v for k, v in vars(__builtins__).items()
                   if k in ("abs", "round", "min", "max", "pow")}
        result = eval(expression, {"__builtins__": allowed})
        return json.dumps({"expression": expression, "result": result})
    except Exception as e:
        return json.dumps({"error": str(e)})


def execute_tool(tool_name: str, tool_input: dict) -> str:
    if tool_name == "get_weather":
        return get_weather(**tool_input)
    elif tool_name == "calculate":
        return calculate(**tool_input)
    else:
        return json.dumps({"error": f"Unknown tool: {tool_name}"})


# ── Agent Class ──────────────────────────────────────────────────────────────

class ClaudeAgent:
    def __init__(self):
        self.client = anthropic.Anthropic()
        self.model = "claude-sonnet-4-6"
        self.messages = []

    def add_message(self, role: str, content):
        self.messages.append({"role": role, "content": content})

    def _call_claude_with_retry(self, max_retries: int = 3) -> anthropic.types.Message:
        for attempt in range(max_retries):
            try:
                return self.client.messages.create(
                    model=self.model,
                    max_tokens=4096,
                    tools=TOOLS,
                    messages=self.messages
                )
            except anthropic.RateLimitError:
                if attempt == max_retries - 1:
                    raise
                wait = 2 ** attempt
                print(f"  Rate limited. Retrying in {wait}s...")
                time.sleep(wait)
            except anthropic.APIStatusError as e:
                if attempt == max_retries - 1:
                    raise
                print(f"  API error {e.status_code}. Retrying...")
                time.sleep(2 ** attempt)

    def _agentic_loop(self) -> str:
        max_iterations = 10
        iteration = 0

        while iteration < max_iterations:
            iteration += 1

            response = self._call_claude_with_retry()
            stop_reason = response.stop_reason

            if stop_reason == "end_turn":
                final_text = next(
                    (block.text for block in response.content if hasattr(block, "text")),
                    "No response generated."
                )
                return final_text

            elif stop_reason == "tool_use":
                self.add_message("assistant", response.content)
                tool_results = []
                for block in response.content:
                    if block.type == "tool_use":
                        print(f"  → Agent calling tool: {block.name}({block.input})")
                        try:
                            result = execute_tool(block.name, block.input)
                        except Exception as e:
                            result = json.dumps({"error": str(e)})
                        tool_results.append({
                            "type": "tool_result",
                            "tool_use_id": block.id,
                            "content": result
                        })
                self.add_message("user", tool_results)

            else:
                return f"Unexpected stop reason: {stop_reason}"

        return "Max iterations reached without a final answer."

    def run(self, user_input: str) -> str:
        self.add_message("user", user_input)
        return self._agentic_loop()


# ── Run It ───────────────────────────────────────────────────────────────────

if __name__ == "__main__":
    agent = ClaudeAgent()

    question = (
        "What's the weather like in Naples, FL right now? "
        "And what is that temperature in Celsius?"
    )

    print(f"User: {question}\n")
    answer = agent.run(question)
    print(f"\nAgent: {answer}")

Run it with python agent.py. Here's what you'll see:

example output
User: What's the weather like in Naples, FL right now? And what is that temperature in Celsius?

  → Agent calling tool: get_weather({'city': 'Naples, FL'})
  → Agent calling tool: calculate({'expression': '(91 - 32) * 5 / 9'})

Agent: It's currently 91°F (32.8°C) and sunny in Naples, FL.

Two tool calls, one clean answer. The agent figured out it needed both tools on its own — you didn't hard-code that logic anywhere.

How It Works

Here's the plain-English version of what just happened. You sent Claude a question. Claude looked at the available tools, decided it needed weather data and a unit conversion, and returned a response with stop_reason: "tool_use" instead of just answering.

Your loop caught that, ran both tools, and fed the results back as a new user message. Claude read those results and this time returned stop_reason: "end_turn" with the final answer. That's the whole Claude API agentic loops tutorial in one paragraph.

The key insight is that the model never runs your code — it just asks for results in a structured format, and you decide how to fulfill those requests. That separation is what makes agents safe and composable. You stay in control of what the tools actually