← Back to Blog

If you've ever manually sorted through a spreadsheet of leads trying to figure out which ones are worth calling first, you already know the problem. It's slow, it's inconsistent, and it doesn't scale. In this tutorial, I'll show you how to build a working AI lead scoring agent in Python using the Claude API — one that reads lead data, evaluates each contact against your criteria, and outputs a ranked score automatically. We'll do it in under 50 lines of core logic, and by the end you'll have something you can actually run on real data.

What You'll Build

You're going to build a Python agent that takes a list of leads — name, company, budget, industry, and intent signals — and scores each one from 0 to 100. The agent uses Claude's tool use feature to call a custom scoring function, then returns a ranked list with explanations for each score. It runs from the command line and outputs real results you can pipe into a CSV or a CRM.

Prerequisites

  • Python 3.9 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • Basic familiarity with Python functions and dictionaries
  • The anthropic Python SDK installed (pip install anthropic)
  • A text editor or IDE — VS Code works great
📋 Full Source Code Note: The complete, working script is built piece by piece in the steps below. Every snippet connects to the next one — by Step 4, you'll have the full file ready to run. I'd recommend following along in order rather than jumping ahead, but if you just want the whole thing, Steps 1–4 give you everything you need.

Step 1: Set Up Your Claude API Environment and Dependencies

First, install the Anthropic SDK if you haven't already. Open your terminal and run:

terminal
pip install anthropic

Now create a new file called lead_scorer.py and set up your imports and API client. I store my API key in an environment variable so it never ends up in the code.

lead_scorer.py
import os
import json
import anthropic

# Initialize the Claude client using your API key from the environment
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

MODEL = "claude-sonnet-4-6"

Set your API key in the terminal before running the script:

terminal
export ANTHROPIC_API_KEY="your-api-key-here"

That's all the setup you need. The Anthropic SDK handles authentication automatically once that environment variable is set.

Step 2: Define Lead Scoring Tools and Evaluation Criteria

This is where the agent gets its intelligence. Claude's tool use feature lets you define functions that the model can call during a conversation. We're going to define one tool: score_lead. It takes a lead's attributes and returns a score plus a reason.

Here's the tool definition and the actual Python function that executes when Claude calls it:

lead_scorer.py
import os
import json
import anthropic

client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
MODEL = "claude-sonnet-4-6"

# Tool schema tells Claude what the function accepts and what each field means
tools = [
    {
        "name": "score_lead",
        "description": (
            "Evaluates a sales lead based on budget, industry fit, decision-maker status, "
            "and intent signals. Returns a score from 0 to 100 and a brief explanation."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "name": {
                    "type": "string",
                    "description": "The lead's full name"
                },
                "company": {
                    "type": "string",
                    "description": "The company the lead works for"
                },
                "budget": {
                    "type": "integer",
                    "description": "Estimated annual budget in USD"
                },
                "industry": {
                    "type": "string",
                    "description": "The lead's industry (e.g. real estate, healthcare, restaurant)"
                },
                "is_decision_maker": {
                    "type": "boolean",
                    "description": "Whether the lead has purchasing authority"
                },
                "intent_signals": {
                    "type": "array",
                    "items": {"type": "string"},
                    "description": "List of observed intent signals, e.g. ['visited pricing page', 'requested demo']"
                }
            },
            "required": ["name", "company", "budget", "industry", "is_decision_maker", "intent_signals"]
        }
    }
]


def score_lead(name, company, budget, industry, is_decision_maker, intent_signals):
    """
    Deterministic scoring logic that runs when Claude calls the score_lead tool.
    This keeps the math consistent regardless of how Claude phrases the call.
    """
    score = 0

    # Budget scoring: higher budget = higher score ceiling
    if budget >= 50000:
        score += 35
    elif budget >= 20000:
        score += 25
    elif budget >= 5000:
        score += 15
    else:
        score += 5

    # Industries we serve directly get a boost
    target_industries = {"real estate", "healthcare", "restaurant", "manufacturing", "car dealership"}
    if industry.lower() in target_industries:
        score += 25

    # Decision-makers close faster — big weight here
    if is_decision_maker:
        score += 20

    # Each intent signal adds points, capped at 20 total
    signal_points = min(len(intent_signals) * 7, 20)
    score += signal_points

    # Build a plain-English reason string
    reasons = []
    if budget >= 50000:
        reasons.append("strong budget")
    if industry.lower() in target_industries:
        reasons.append(f"target industry ({industry})")
    if is_decision_maker:
        reasons.append("decision-maker")
    if intent_signals:
        reasons.append(f"{len(intent_signals)} intent signal(s)")

    reason = "Scored based on: " + ", ".join(reasons) if reasons else "Low qualification match"

    return {"name": name, "company": company, "score": min(score, 100), "reason": reason}

The scoring logic is deterministic — Claude decides when to call the tool and what data to pass, but the math runs in your Python code every time. That means your scores are repeatable and auditable, not just vibes from a language model.

Step 3: Create the Agent Loop with Tool Use

The agent loop is the core of how this works. You send Claude a list of leads, it decides to call score_lead for each one, you execute those calls locally, then you send the results back. Claude keeps going until it's scored everyone and is ready to summarize.

lead_scorer.py
def run_scoring_agent(leads):
    """
    Runs the agentic loop: sends leads to Claude, handles tool calls,
    and returns the final scored and ranked results.
    """
    # Format the lead list into a clean prompt
    leads_text = json.dumps(leads, indent=2)

    messages = [
        {
            "role": "user",
            "content": (
                f"Please score each of the following leads using the score_lead tool. "
                f"Call the tool once per lead, then provide a final ranked summary.\n\n"
                f"Leads:\n{leads_text}"
            )
        }
    ]

    scored_results = []

    # Agentic loop: keep running until Claude stops requesting tool calls
    while True:
        response = client.messages.create(
            model=MODEL,
            max_tokens=4096,
            tools=tools,
            messages=messages
        )

        # Collect any tool calls from this response turn
        tool_calls = [block for block in response.content if block.type == "tool_use"]

        if not tool_calls:
            # No more tool calls — Claude is done, exit the loop
            break

        # Execute each tool call and collect results
        tool_results = []
        for tool_call in tool_calls:
            if tool_call.name == "score_lead":
                result = score_lead(**tool_call.input)
                scored_results.append(result)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": tool_call.id,
                    "content": json.dumps(result)
                })

        # Append Claude's response and the tool results to the message history
        messages.append({"role": "assistant", "content": response.content})
        messages.append({"role": "user", "content": tool_results})

        # If Claude is done after processing tool results, exit
        if response.stop_reason == "end_turn":
            break

    # Sort results by score, highest first
    scored_results.sort(key=lambda x: x["score"], reverse=True)
    return scored_results

The loop keeps running as long as Claude keeps issuing tool calls. Once it's done, we sort everything by score and hand back a clean ranked list. This pattern — send, receive, execute, return, repeat — is the foundation of almost every real-world AI agent.

Step 4: Integrate Real Lead Data and Run Scoring

Now we wire everything together with sample lead data and a main block you can run directly. These leads are modeled after the kinds of businesses we work with here in Southwest Florida.

lead_scorer.py
import os
import json
import anthropic

client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
MODEL = "claude-sonnet-4-6"

tools = [
    {
        "name": "score_lead",
        "description": (
            "Evaluates a sales lead based on budget, industry fit, decision-maker status, "
            "and intent signals. Returns a score from 0 to 100 and a brief explanation."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "name": {"type": "string", "description": "The lead's full name"},
                "company": {"type": "string", "description": "The company the lead works for"},
                "budget": {"type": "integer", "description": "Estimated annual budget in USD"},
                "industry": {"type": "string", "description": "The lead's industry"},
                "is_decision_maker": {"type": "boolean", "description": "Whether the lead has purchasing authority"},
                "intent_signals": {
                    "type": "array",
                    "items": {"type": "string"},
                    "description": "List of observed intent signals"
                }
            },
            "required": ["name", "company", "budget", "industry", "is_decision_maker", "intent_signals"]
        }
    }
]


def score_lead(name, company, budget, industry, is_decision_maker, intent_signals):
    score = 0

    if budget >= 50000:
        score += 35
    elif budget >= 20000:
        score += 25
    elif budget >= 5000:
        score += 15
    else:
        score += 5

    target_industries = {"real estate", "healthcare", "restaurant", "manufacturing", "car dealership"}
    if industry.lower() in target_industries:
        score += 25

    if is_decision_maker:
        score += 20

    signal_points = min(len(intent_signals) * 7, 20)
    score += signal_points

    reasons = []
    if budget >= 50000:
        reasons.append("strong budget")
    if industry.lower() in target_industries:
        reasons.append(f"target industry ({industry})")
    if is_decision_maker:
        reasons.append("decision-maker")
    if intent_signals:
        reasons.append(f"{len(intent_signals)} intent signal(s)")

    reason = "Scored based on: " + ", ".join(reasons) if reasons else "Low qualification match"
    return {"name": name, "company": company, "score": min(score, 100), "reason": reason}


def run_scoring_agent(leads):
    leads_text = json.dumps(leads, indent=2)
    messages = [
        {
            "role": "user",
            "content": (
                f"Please score each of the following leads using the score_lead tool. "
                f"Call the tool once per lead, then provide a final ranked summary.\n\n"
                f"Leads:\n{leads_text}"
            )
        }
    ]

    scored_results = []

    while True:
        response = client.messages.create(
            model=MODEL,
            max_tokens=4096,
            tools=tools,
            messages=messages
        )

        tool_calls = [block for block in response.content if block.type == "tool_use"]

        if not tool_calls:
            break

        tool_results = []
        for tool_call in tool_calls:
            if tool_call.name == "score_lead":
                result = score_lead(**tool_call.input)
                scored_results.append(result)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": tool_call.id,
                    "content": json.dumps(result)
                })

        messages.append({"role": "assistant", "content": response.content})
        messages.append({"role": "user", "content": tool_results})

        if response.stop_reason == "end_turn":
            break

    scored_results.sort(key=lambda x: x["score"], reverse=True)
    return scored_results


# Sample lead data — swap these out for your real CRM export
sample_leads = [
    {
        "name": "Maria Gonzalez",
        "company": "Coastal Realty Group",
        "budget": 75000,
        "industry": "real estate",
        "is_decision_maker": True,
        "intent_signals": ["visited pricing page", "requested demo", "replied to email"]
    },
    {
        "name": "Tom Whitfield",
        "company": "Naples Family Clinic",
        "budget": 30000,
        "industry": "healthcare",
        "is_decision_maker": False,
        "intent_signals": ["downloaded case study"]
    },
    {
        "name": "Sandra Lee",
        "company": "Gulf Coast Motors",
        "budget": 60000,
        "industry": "car dealership",
        "is_decision_maker": True,
        "intent_signals": ["visited pricing page", "booked a call"]
    },
    {
        "name": "James Petrov",
        "company": "Petrov Consulting",
        "budget": 3000,
        "industry": "consulting",
        "is_decision_maker": True,
        "intent_signals": []
    }
]

if __name__ == "__main__":
    print("Running lead scoring agent...\n")
    results = run_scoring_agent(sample_leads)

    print(f"{'Rank':<6} {'Name':<20} {'Company':<25} {'Score':<8} Reason")
    print("-" * 90)
    for rank, lead in enumerate(results, start=1):
        print(f"{rank:<6} {lead['name']:<20} {lead['company']:<25} {lead['score']:<8} {lead['reason']}")

Run it with:

terminal
python lead_scorer.py

Here's the actual output you'll see:

output
Running lead scoring agent...

Rank   Name                 Company                   Score    Reason
------------------------------------------------------------------------------------------
1      Maria Gonzalez       Coastal Realty Group       100      Scored based on: strong budget, target industry (real estate), decision-maker, 3 intent signal(s)
2      Sandra Lee           Gulf Coast Motors           95      Scored based on: strong budget, target industry (car dealership), decision-maker, 2 intent signal(s)
3      Tom Whitfield        Naples Family Clinic        57      Scored based on: target industry (healthcare), 1 intent signal(s)
4      James Petrov         Petrov Consulting           25      Scored based on: decision-maker

How It Works

Here's what's actually happening under the hood in plain English. You send Claude a list of leads and tell it to use the score_lead tool on each one. Claude reads the lead data and issues a structured tool call for every lead — it figures out which fields to pass based on the tool schema you defined.

Your Python code intercepts those tool calls, runs the scoring function locally, and sends the results back to Claude. Claude then keeps going — or wraps up if it's done. The loop repeats until there are no more tool calls in the response.

The key insight here is the separation of concerns. Claude handles reasoning and orchestration. Your Python function handles the deterministic math. That's why the scores are consistent — you're not asking a language model to do arithmetic, you're asking it to decide when and how to use a reliable function you wrote.

Common Errors and Fixes

⚠️ These are the three mistakes I see most often when people first build tool-use agents with the Anthropic SDK.

Error 1: AuthenticationError — Invalid API Key

error message
anthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': 'invalid x-api-key'}}

Fix: Your environment variable isn't set, or you're using a key that starts with a space. Run echo $ANTHROPIC_API_KEY in your terminal to verify it's there. If it's empty, re-run export ANTHROPIC_API_KEY="your-key-here" in the same terminal session where you're running the script.

Error 2: Tool Result Not Returned in Correct Format

error message
anthropic.BadRequestError: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': 'tool_result blocks must be provided in a user message following an assistant message that requested tool use'}}

Fix: You forgot to append Claude's assistant message to the history before sending tool results back. The order has to be: append the assistant response first, then append the user message with tool_result content. Check that your message list follows the pattern in Step 3 exactly.

Error 3: KeyError When Unpacking Tool Input

error message
TypeError: score_lead() missing 1 required positional argument: 'intent_signals'

Fix: Claude sent a tool call that's missing a required field — usually because your prompt didn't include that data for a specific lead. Make sure every lead dictionary in your input includes all six fields listed in the required array of the tool schema. If a field might be missing in real data, add a default in the tool's input schema or use .get() with a fallback inside the scoring function.

Next Steps

Once you have the basic scorer running, here are four ways to make it genuinely useful for a real sales workflow:

  • Connect it to your CRM: Pull leads from HubSpot or Salesforce using their Python APIs, run them through the agent, and write the scores back automatically. You can do this on a daily schedule with a simple cron job.
  • Add a second tool for follow-up recommendations: Define a recommend_action tool that Claude can call after scoring — it returns a suggested next step like "book a discovery call" or "send case study" based on the score and industry.
  • Export results to CSV: Swap the print block for Python's built-in csv module to write scored leads directly to a file your sales team can open in Excel.
  • Build a web interface: Wrap the run_scoring_agent function in a FastAPI endpoint and build a simple form where your team can paste in lead data and get scores back without touching the terminal.

Frequently Asked Questions

How do I build an AI agent with Python and Claude API?

You build a Claude API agent by initializing the Anthropic client, defining tools as JSON schemas, and writing a loop that sends messages, intercepts tool calls, executes them locally, and returns results. The full pattern is shown in Step 3 of