← Back to Blog

What You'll Build

You're going to build a production-ready lead scoring agent in Python that uses Claude's tool-use feature to evaluate incoming leads and rank them by sales priority. It reads lead data, reasons through a set of scoring criteria, and returns a structured score with a written explanation — automatically.

By the end of this tutorial you'll have a working agentic loop that handles multi-turn conversations, calls custom tools, and outputs a score from 0–100 for any lead you throw at it. I'll use real estate leads as the example since that's a vertical we work with constantly at Naples AI, but the same pattern works for any industry.

📦 Full Source Code
All the code in this tutorial is complete and runnable — no pseudocode, no placeholders. Each step builds on the last, and by Step 5 you'll have the entire working agent. Copy each block in order and you're good to go.

Prerequisites

  • Python 3.10 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • Basic familiarity with Python classes and dictionaries
  • anthropic Python SDK installed (pip install anthropic)
  • python-dotenv for managing your API key (pip install python-dotenv)

Step 1: Set Up Your Claude API Environment and Project Structure

Start by creating a clean project folder and setting up your environment file. You never want to hardcode your API key directly in your script — that's how keys get accidentally committed to GitHub.

Here's the folder structure we're working with:

lead-scoring-agent/
├── .env
├── agent.py
├── tools.py
└── run.py

Create your .env file first:

.env
ANTHROPIC_API_KEY=your_api_key_here

Now install your dependencies if you haven't already:

pip install anthropic python-dotenv

Let's verify the SDK connects correctly before writing any agent logic. This quick sanity check saves you from chasing errors later.

test_connection.py
import os
from dotenv import load_dotenv
import anthropic

load_dotenv()

client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=64,
    messages=[{"role": "user", "content": "Say: connection successful"}]
)

print(response.content[0].text)

Run it with python test_connection.py. If you see "connection successful" printed back, you're ready to build.

Step 2: Define Lead Scoring Criteria and Tool Definitions

This is where Claude gets its "job description." We define two tools: one to score a lead based on specific criteria, and one to flag a lead for immediate follow-up. These tools are passed to Claude as JSON schemas, and Claude decides when and how to call them.

The scoring criteria I'm using here are built around real estate leads — budget, timeline, pre-approval status, property type, and engagement level. You can swap these out for whatever matters in your sales process.

tools.py
import os
from dotenv import load_dotenv

load_dotenv()

# Tool definitions passed to Claude as JSON schema
LEAD_SCORING_TOOLS = [
    {
        "name": "score_lead",
        "description": (
            "Evaluate a real estate lead and assign a score from 0 to 100 based on "
            "the provided criteria. Higher scores mean higher sales priority. "
            "Call this tool once you have enough information to evaluate the lead."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "lead_score": {
                    "type": "integer",
                    "description": "Numeric score from 0 to 100 reflecting lead quality.",
                    "minimum": 0,
                    "maximum": 100
                },
                "priority_tier": {
                    "type": "string",
                    "enum": ["Hot", "Warm", "Cold"],
                    "description": "Priority classification based on score: Hot (70-100), Warm (40-69), Cold (0-39)."
                },
                "reasoning": {
                    "type": "string",
                    "description": "Plain-English explanation of why this score was assigned, referencing specific lead attributes."
                },
                "recommended_action": {
                    "type": "string",
                    "description": "Specific next step for the sales team to take with this lead."
                },
                "score_breakdown": {
                    "type": "object",
                    "description": "Individual scores for each evaluated criterion.",
                    "properties": {
                        "budget_fit": {"type": "integer", "minimum": 0, "maximum": 25},
                        "timeline_urgency": {"type": "integer", "minimum": 0, "maximum": 25},
                        "financial_readiness": {"type": "integer", "minimum": 0, "maximum": 25},
                        "engagement_level": {"type": "integer", "minimum": 0, "maximum": 25}
                    },
                    "required": ["budget_fit", "timeline_urgency", "financial_readiness", "engagement_level"]
                }
            },
            "required": ["lead_score", "priority_tier", "reasoning", "recommended_action", "score_breakdown"]
        }
    },
    {
        "name": "flag_for_immediate_followup",
        "description": (
            "Flag a lead for immediate human follow-up when they show signals of being "
            "ready to transact within 30 days or have expressed urgency. "
            "Use this in addition to score_lead for hot leads."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "flag_reason": {
                    "type": "string",
                    "description": "Why this lead needs immediate attention from a human agent."
                },
                "contact_method": {
                    "type": "string",
                    "enum": ["phone", "email", "text"],
                    "description": "Recommended contact method based on lead preferences."
                },
                "urgency_signals": {
                    "type": "array",
                    "items": {"type": "string"},
                    "description": "List of specific signals that triggered this flag."
                }
            },
            "required": ["flag_reason", "contact_method", "urgency_signals"]
        }
    }
]
💡 Why two tools?
Having a separate flagging tool lets Claude take two distinct actions on the same lead. A hot lead might trigger both score_lead and flag_for_immediate_followup in a single turn. This is the composable, multi-action pattern that makes agentic loops powerful.

Step 3: Build the Main Agent Class with Tool Use

Now we wire everything together in the agent class. This class holds the conversation history, sends messages to Claude, and handles tool calls when Claude decides to use them. The process_lead method is your main entry point.

agent.py
import os
import json
from dotenv import load_dotenv
import anthropic
from tools import LEAD_SCORING_TOOLS

load_dotenv()

SYSTEM_PROMPT = """You are an expert real estate sales qualification agent working for a Southwest Florida real estate brokerage.

Your job is to evaluate incoming leads and assign them a score from 0-100 based on these criteria:

SCORING CRITERIA (25 points each):
1. Budget Fit (0-25): How well does their budget align with available inventory?
   - 25: Budget matches or exceeds market for desired property type
   - 15: Budget is slightly below market but workable
   - 5: Budget is significantly below market

2. Timeline Urgency (0-25): How soon do they need to buy or sell?
   - 25: Within 30 days
   - 15: Within 90 days
   - 5: 6+ months or undefined

3. Financial Readiness (0-25): Are they pre-approved or paying cash?
   - 25: Cash buyer or fully pre-approved
   - 15: Pre-approval in progress
   - 5: No pre-approval, just browsing

4. Engagement Level (0-25): How responsive and serious are they?
   - 25: Requested specific showing, asked detailed questions
   - 15: Filled out contact form with some details
   - 5: Generic inquiry, minimal information

Always call the score_lead tool after your evaluation. For leads scoring 70+, also call flag_for_immediate_followup.
Be direct, specific, and base your reasoning only on the data provided."""


class LeadScoringAgent:
    def __init__(self):
        self.client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
        self.model = "claude-sonnet-4-6"
        self.conversation_history = []
        self.tool_results = {}

    def _handle_tool_call(self, tool_name: str, tool_input: dict) -> str:
        """Process a tool call and return the result as a string."""
        # Store results so we can return them to the caller
        self.tool_results[tool_name] = tool_input

        if tool_name == "score_lead":
            result = {
                "status": "scored",
                "lead_score": tool_input["lead_score"],
                "priority_tier": tool_input["priority_tier"]
            }
        elif tool_name == "flag_for_immediate_followup":
            result = {
                "status": "flagged",
                "flag_confirmed": True,
                "contact_method": tool_input["contact_method"]
            }
        else:
            result = {"status": "error", "message": f"Unknown tool: {tool_name}"}

        return json.dumps(result)

    def process_lead(self, lead_data: dict) -> dict:
        """
        Takes a lead dictionary and runs it through the scoring agent.
        Returns the complete scoring result.
        """
        # Reset state for each new lead
        self.conversation_history = []
        self.tool_results = {}

        # Format lead data as the initial user message
        lead_message = f"""Please evaluate this real estate lead and score them:

Lead Information:
{json.dumps(lead_data, indent=2)}

Analyze this lead against the scoring criteria and call the appropriate tools."""

        self.conversation_history.append({
            "role": "user",
            "content": lead_message
        })

        # Run the agentic loop
        final_result = self._run_agentic_loop()
        return final_result

    def _run_agentic_loop(self) -> dict:
        """
        The core agentic loop. Keeps running until Claude stops calling tools.
        Returns the collected tool results.
        """
        while True:
            response = self.client.messages.create(
                model=self.model,
                max_tokens=4096,
                system=SYSTEM_PROMPT,
                tools=LEAD_SCORING_TOOLS,
                messages=self.conversation_history
            )

            # Add Claude's response to history
            self.conversation_history.append({
                "role": "assistant",
                "content": response.content
            })

            # If Claude is done (no more tool calls), exit the loop
            if response.stop_reason == "end_turn":
                break

            # If Claude wants to use tools, process each one
            if response.stop_reason == "tool_use":
                tool_results_for_history = []

                for block in response.content:
                    if block.type == "tool_use":
                        tool_result = self._handle_tool_call(block.name, block.input)

                        tool_results_for_history.append({
                            "type": "tool_result",
                            "tool_use_id": block.id,
                            "content": tool_result
                        })

                # Return tool results to Claude so it can continue
                self.conversation_history.append({
                    "role": "user",
                    "content": tool_results_for_history
                })

        return self.tool_results

Step 4: Implement the Agentic Loop with Conversation History

The _run_agentic_loop method above is the heart of the agent. Let me explain what's actually happening in that while True block, because this pattern shows up in every serious AI agent you'll build.

Claude doesn't execute your tools — it just tells you which tool to call and with what arguments. Your Python code runs the actual tool logic, then sends the result back to Claude as a new message. Claude reads that result and decides what to do next.

The loop exits when response.stop_reason == "end_turn", which means Claude has finished reasoning and doesn't need to call any more tools. Now let's build the run script that ties everything together.

run.py
import os
import json
from dotenv import load_dotenv
from agent import LeadScoringAgent

load_dotenv()

def format_score_output(lead_name: str, result: dict) -> None:
    """Print a formatted score report to the console."""
    print("\n" + "="*60)
    print(f"LEAD SCORING REPORT: {lead_name}")
    print("="*60)

    if "score_lead" in result:
        score_data = result["score_lead"]
        print(f"\n📊 OVERALL SCORE:    {score_data['lead_score']}/100")
        print(f"🏷️  PRIORITY TIER:   {score_data['priority_tier']}")
        print(f"\n📋 SCORE BREAKDOWN:")

        breakdown = score_data.get("score_breakdown", {})
        print(f"   Budget Fit:          {breakdown.get('budget_fit', 0)}/25")
        print(f"   Timeline Urgency:    {breakdown.get('timeline_urgency', 0)}/25")
        print(f"   Financial Readiness: {breakdown.get('financial_readiness', 0)}/25")
        print(f"   Engagement Level:    {breakdown.get('engagement_level', 0)}/25")

        print(f"\n🧠 REASONING:")
        print(f"   {score_data['reasoning']}")

        print(f"\n✅ RECOMMENDED ACTION:")
        print(f"   {score_data['recommended_action']}")

    if "flag_for_immediate_followup" in result:
        flag_data = result["flag_for_immediate_followup"]
        print(f"\n🚨 IMMEDIATE FOLLOW-UP FLAGGED")
        print(f"   Contact via:  {flag_data['contact_method'].upper()}")
        print(f"   Reason:       {flag_data['flag_reason']}")
        print(f"   Signals:      {', '.join(flag_data['urgency_signals'])}")

    print("\n" + "="*60)


def main():
    agent = LeadScoringAgent()

    # Sample leads for testing — modeled after real Southwest Florida real estate leads
    sample_leads = [
        {
            "name": "Marcus & Jennifer Delgado",
            "source": "Zillow Inquiry",
            "property_interest": "Single-family home, Naples FL",
            "budget": "$850,000 - $1,100,000",
            "timeline": "Want to be in before January, moving from Chicago",
            "financing": "Pre-approved with Fifth Third Bank, 20% down",
            "notes": "Requested 3 specific showings this week, asked about HOA fees and flood zone ratings. Very responsive to emails.",
            "engagement": "Called office twice, toured 2 properties already"
        },
        {
            "name": "Robert Thorne",
            "source": "Website Contact Form",
            "property_interest": "Condo, downtown Naples or Bonita Springs",
            "budget": "Around $300k",
            "timeline": "Sometime next year maybe",
            "financing": "Not sure yet, haven't talked to a bank",
            "notes": "Generic message: 'interested in condos in the area'",
            "engagement": "Submitted one form, hasn't responded to follow-up email"
        }
    ]

    for lead in sample_leads:
        print(f"\nProcessing lead: {lead['name']}...")
        result = agent.process_lead(lead)
        format_score_output(lead["name"], result)


if __name__ == "__main__":
    main()

Step 5: Test with Sample Lead Data and Verify Scoring Output

Run the agent with python run.py. You should see two complete scoring reports printed to your terminal. Here's what the actual output looks like for the two sample leads:

Sample Output
Processing lead: Marcus & Jennifer Delgado...

============================================================
LEAD SCORING REPORT: Marcus & Jennifer Delgado
============================================================

📊 OVERALL SCORE:    92/100
🏷️  PRIORITY TIER:   Hot

📋 SCORE BREAKDOWN:
   Budget Fit:          23/25
   Timeline Urgency:    25/25
   Financial Readiness: 24/25
   Engagement Level:    20/25

🧠 REASONING:
   Marcus and Jennifer are an exceptionally strong lead. Their $850K-$1.1M budget
   aligns well with Naples single-family inventory in that range. Their January
   move-in deadline creates genuine urgency — they're relocating from Chicago and
   need to close within roughly 60-90 days. Full pre-approval with 20% down
   removes financing risk almost entirely. They've already toured properties,
   called the office twice, and asked specific due-diligence questions about HOA
   and flood zones — clear signals of a serious, informed buyer.

✅ RECOMMENDED ACTION:
   Call Marcus or Jennifer directly today. Prepare a shortlist of 4-5 properties
   in the $900K-$1.05M range that cleared their flood zone concern. Offer a
   dedicated showing day this week while their travel schedule allows.

🚨 IMMEDIATE FOLLOW-UP FLAGGED
   Contact via:  PHONE
   Reason:       Buyer has a hard relocation deadline before January and is
                 already actively touring. High risk of closing with a competitor
                 agent if not engaged within 24 hours.
   Signals:      Pre-approved buyer, active touring, relocation deadline,
                 repeated office contact

============================================================

Processing lead: Robert Thorne...

============================================================
LEAD SCORING REPORT: Robert Thorne
============================================================

📊 OVERALL SCORE:    18/100
🏷️  PRIORITY TIER:   Cold

📋 SCORE BREAKDOWN:
   Budget Fit:          8/25
   Timeline Urgency:    2/25
   Financial Readiness: 3/25
   Engagement Level:    5/25

🧠 REASONING:
   Robert's inquiry shows early-stage browsing behavior with no commitment signals.
   His $300K budget is significantly below the Naples condo market median, limiting
   available inventory. His timeline of "sometime next year maybe" indicates no
   urgency. He has not begun pre-approval conversations. He submitted one generic
   form message and has not responded to follow-up — low engagement overall.

✅ RECOMMENDED ACTION:
   Add Robert to a long-term email nurture sequence. Send a Naples condo market
   report to build credibility. Follow up in 30 days with a check-in email.
   Do not prioritize phone outreach at this stage.

============================================================

That's a real scoring output from Claude — not mocked data. The reasoning is specific to each lead's actual attributes, and the recommended actions are genuinely different based on the score.

How It Works: Multi-Turn Conversations, Tool Invocation, and Reasoning

Claude doesn't just classify a lead into a bucket. It reasons through each criterion, weighs the evidence, and then decides which tool to call with what arguments. That reasoning is what separates this from a simple if/else scoring script.

Here's the flow in plain English:

  1. You send lead data as a user message with the tool definitions attached.
  2. Claude reads the lead, reasons through the scoring criteria in its context window, and generates a tool call with a fully populated JSON input.
  3. Your Python code receives that JSON, runs _handle_tool_call, and sends the result back to Claude as a tool result message.
  4. Claude reads the confirmation and decides if it needs to call another tool (like flag_for_immediate_followup for hot leads).
  5. When it's done calling tools, it returns an end_turn stop reason and the loop exits.

The conversation history is the key piece. Every message — your inputs, Claude's responses, tool calls, and tool results — gets appended to self.conversation_history. This gives Claude full context at every step so it doesn't lose track of what it already decided.

🔁 Why not just parse Claude's text response?
You could ask Claude to output JSON in its text response and parse it yourself — but tool use is more reliable. Claude is trained to produce well-