← Back to Blog

What You'll Build

If you've ever wasted time chasing leads that went nowhere, this tutorial is for you. You're going to build a production-ready AI lead scoring agent in Python that uses Claude's tool-use capability to evaluate incoming leads, assign a priority score from 1–100, and explain its reasoning in plain English. By the end, you'll have a working agent loop you can connect to any CRM or intake form.

Prerequisites

  • Python 3.10 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • anthropic Python SDK installed: pip install anthropic
  • Basic familiarity with Python functions and classes
  • Optional: a CRM or webhook endpoint to route output to
📦 Full Source Code
All the code in this tutorial is production-ready and syntactically complete. Each step below builds on the last, and by Step 4 you'll have the entire working agent. Copy the snippets in order, or scroll to Step 3 to see the complete agent class assembled in one place.

Step 1: Initialize the Claude Client and Define Scoring Criteria

The first thing we need is a clean client setup and a set of scoring criteria that Claude will use as its rubric. Think of this as writing the instructions you'd give a sharp sales development rep on their first day.

We're using claude-sonnet-4-6 here because it handles multi-tool reasoning well without burning through your budget on every lead evaluation. The system prompt is where you encode your business logic — what makes a lead worth pursuing.

client_setup.py
import os
import anthropic

# Initialize the Anthropic client using your API key
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))

MODEL = "claude-sonnet-4-6"

# This system prompt is your scoring rubric — edit it to match your business
SYSTEM_PROMPT = """
You are an expert lead qualification agent for a B2B software company.
Your job is to evaluate incoming leads and assign a score from 1 to 100 based on these criteria:

SCORING CRITERIA:
- Budget fit (does their stated budget match our pricing?): 0–25 points
- Company size (employee count signals deal size): 0–20 points
- Decision-maker authority (are they the buyer?): 0–20 points
- Timeline urgency (how soon do they need a solution?): 0–20 points
- Industry fit (do we have case studies in their vertical?): 0–15 points

SCORE TIERS:
- 80–100: Hot lead — prioritize immediately
- 60–79:  Warm lead — follow up within 24 hours
- 40–59:  Nurture — add to drip sequence
- 0–39:   Disqualify — do not spend sales time here

Always use the available tools to break your evaluation into clear steps.
Explain your reasoning before giving a final score.
"""
      

Step 2: Create Tool Functions for Lead Evaluation

This is where Claude gets its superpowers. Instead of asking the model to score a lead in one shot, we give it a set of tools it can call — one per scoring dimension. This keeps the reasoning modular and auditable, which matters a lot when a sales manager asks "why did we deprioritize this account?"

Each tool is a plain Python function paired with a JSON schema that tells Claude what arguments to pass. The schema is what Claude reads; the function is what actually runs when Claude decides to use the tool.

tools.py
import json

# --- Tool Definitions (what Claude sees) ---

TOOLS = [
    {
        "name": "score_budget_fit",
        "description": "Evaluate how well the lead's stated budget aligns with our pricing. Returns a score from 0 to 25.",
        "input_schema": {
            "type": "object",
            "properties": {
                "stated_budget": {
                    "type": "string",
                    "description": "The lead's stated or implied budget range, e.g. '$5,000/month' or 'under $10k'"
                },
                "reasoning": {
                    "type": "string",
                    "description": "Brief explanation of why you assigned this score"
                },
                "score": {
                    "type": "integer",
                    "description": "Budget fit score from 0 to 25"
                }
            },
            "required": ["stated_budget", "reasoning", "score"]
        }
    },
    {
        "name": "score_company_size",
        "description": "Evaluate the lead's company size as a signal of deal potential. Returns a score from 0 to 20.",
        "input_schema": {
            "type": "object",
            "properties": {
                "employee_count": {
                    "type": "string",
                    "description": "Estimated or stated number of employees"
                },
                "reasoning": {
                    "type": "string",
                    "description": "Brief explanation of why you assigned this score"
                },
                "score": {
                    "type": "integer",
                    "description": "Company size score from 0 to 20"
                }
            },
            "required": ["employee_count", "reasoning", "score"]
        }
    },
    {
        "name": "score_decision_authority",
        "description": "Evaluate whether the lead has buying authority. Returns a score from 0 to 20.",
        "input_schema": {
            "type": "object",
            "properties": {
                "job_title": {
                    "type": "string",
                    "description": "The lead's job title or role"
                },
                "reasoning": {
                    "type": "string",
                    "description": "Brief explanation of why you assigned this score"
                },
                "score": {
                    "type": "integer",
                    "description": "Decision authority score from 0 to 20"
                }
            },
            "required": ["job_title", "reasoning", "score"]
        }
    },
    {
        "name": "score_timeline_urgency",
        "description": "Evaluate how urgently the lead needs a solution. Returns a score from 0 to 20.",
        "input_schema": {
            "type": "object",
            "properties": {
                "stated_timeline": {
                    "type": "string",
                    "description": "The lead's stated implementation timeline, e.g. 'ASAP', 'Q1 next year', 'just exploring'"
                },
                "reasoning": {
                    "type": "string",
                    "description": "Brief explanation of why you assigned this score"
                },
                "score": {
                    "type": "integer",
                    "description": "Timeline urgency score from 0 to 20"
                }
            },
            "required": ["stated_timeline", "reasoning", "score"]
        }
    },
    {
        "name": "score_industry_fit",
        "description": "Evaluate how well the lead's industry matches our target verticals. Returns a score from 0 to 15.",
        "input_schema": {
            "type": "object",
            "properties": {
                "industry": {
                    "type": "string",
                    "description": "The lead's industry or vertical"
                },
                "reasoning": {
                    "type": "string",
                    "description": "Brief explanation of why you assigned this score"
                },
                "score": {
                    "type": "integer",
                    "description": "Industry fit score from 0 to 15"
                }
            },
            "required": ["industry", "reasoning", "score"]
        }
    },
    {
        "name": "compile_final_score",
        "description": "Compile all sub-scores into a final lead score and recommendation. Call this after all other scoring tools.",
        "input_schema": {
            "type": "object",
            "properties": {
                "budget_score": {"type": "integer"},
                "company_size_score": {"type": "integer"},
                "decision_authority_score": {"type": "integer"},
                "timeline_score": {"type": "integer"},
                "industry_score": {"type": "integer"},
                "final_score": {
                    "type": "integer",
                    "description": "Sum of all sub-scores, 0–100"
                },
                "tier": {
                    "type": "string",
                    "enum": ["Hot", "Warm", "Nurture", "Disqualify"]
                },
                "recommended_action": {
                    "type": "string",
                    "description": "Specific next action for the sales team"
                },
                "summary": {
                    "type": "string",
                    "description": "Two-sentence plain-English summary of the lead"
                }
            },
            "required": [
                "budget_score", "company_size_score", "decision_authority_score",
                "timeline_score", "industry_score", "final_score", "tier",
                "recommended_action", "summary"
            ]
        }
    }
]


# --- Tool Implementations (what actually runs) ---

def score_budget_fit(stated_budget: str, reasoning: str, score: int) -> dict:
    return {"dimension": "budget_fit", "stated_budget": stated_budget,
            "reasoning": reasoning, "score": score}

def score_company_size(employee_count: str, reasoning: str, score: int) -> dict:
    return {"dimension": "company_size", "employee_count": employee_count,
            "reasoning": reasoning, "score": score}

def score_decision_authority(job_title: str, reasoning: str, score: int) -> dict:
    return {"dimension": "decision_authority", "job_title": job_title,
            "reasoning": reasoning, "score": score}

def score_timeline_urgency(stated_timeline: str, reasoning: str, score: int) -> dict:
    return {"dimension": "timeline_urgency", "stated_timeline": stated_timeline,
            "reasoning": reasoning, "score": score}

def score_industry_fit(industry: str, reasoning: str, score: int) -> dict:
    return {"dimension": "industry_fit", "industry": industry,
            "reasoning": reasoning, "score": score}

def compile_final_score(budget_score: int, company_size_score: int,
                        decision_authority_score: int, timeline_score: int,
                        industry_score: int, final_score: int, tier: str,
                        recommended_action: str, summary: str) -> dict:
    return {
        "final_score": final_score,
        "tier": tier,
        "breakdown": {
            "budget": budget_score,
            "company_size": company_size_score,
            "decision_authority": decision_authority_score,
            "timeline": timeline_score,
            "industry": industry_score
        },
        "recommended_action": recommended_action,
        "summary": summary
    }

# Map tool names to their Python functions for the agent loop
TOOL_REGISTRY = {
    "score_budget_fit": score_budget_fit,
    "score_company_size": score_company_size,
    "score_decision_authority": score_decision_authority,
    "score_timeline_urgency": score_timeline_urgency,
    "score_industry_fit": score_industry_fit,
    "compile_final_score": compile_final_score
}
      

Step 3: Build the Agent Loop with Tool Use

This is the core of how to build AI agents — the agentic loop. Claude will look at the lead, decide which tools to call and in what order, receive the results, and keep going until it's satisfied it has everything it needs. We just keep the conversation going until Claude signals it's done.

The pattern here is: send message → check if Claude wants to use a tool → run the tool → feed the result back → repeat. It sounds simple because it is, and that simplicity is actually the point.

agent.py
import os
import json
import anthropic

# Import everything we built in the previous steps
from tools import TOOLS, TOOL_REGISTRY
from client_setup import client, MODEL, SYSTEM_PROMPT


class LeadScoringAgent:
    """
    An agentic loop that scores sales leads using Claude and a set of evaluation tools.
    """

    def __init__(self):
        self.client = client
        self.model = MODEL
        self.tools = TOOLS
        self.system_prompt = SYSTEM_PROMPT

    def _run_tool(self, tool_name: str, tool_input: dict) -> str:
        """Look up and execute the requested tool, return result as JSON string."""
        if tool_name not in TOOL_REGISTRY:
            return json.dumps({"error": f"Unknown tool: {tool_name}"})
        tool_fn = TOOL_REGISTRY[tool_name]
        result = tool_fn(**tool_input)
        return json.dumps(result)

    def score_lead(self, lead: dict) -> dict:
        """
        Run the full agentic scoring loop for a single lead.
        Returns the compiled score dict when Claude calls compile_final_score.
        """

        # Format the lead as a clear message for the model
        lead_message = f"""
Please evaluate this incoming lead and score it using the available tools.
Work through each scoring dimension before calling compile_final_score.

LEAD INFORMATION:
Name:      {lead.get('name', 'Unknown')}
Title:     {lead.get('title', 'Unknown')}
Company:   {lead.get('company', 'Unknown')}
Industry:  {lead.get('industry', 'Unknown')}
Employees: {lead.get('employees', 'Unknown')}
Budget:    {lead.get('budget', 'Not stated')}
Timeline:  {lead.get('timeline', 'Not stated')}
Notes:     {lead.get('notes', 'None')}
"""

        messages = [{"role": "user", "content": lead_message}]
        final_result = None

        # Agentic loop — keep going until Claude stops requesting tools
        while True:
            response = self.client.messages.create(
                model=self.model,
                max_tokens=4096,
                system=self.system_prompt,
                tools=self.tools,
                messages=messages
            )

            # Collect any tool calls from this response turn
            tool_calls_made = []
            for block in response.content:
                if block.type == "tool_use":
                    tool_calls_made.append(block)

            # If Claude made no tool calls, we're done
            if not tool_calls_made:
                break

            # Append Claude's full response to the conversation
            messages.append({"role": "assistant", "content": response.content})

            # Execute each tool Claude requested and collect results
            tool_results = []
            for tool_call in tool_calls_made:
                tool_output = self._run_tool(tool_call.name, tool_call.input)

                # Capture the final score when compile_final_score is called
                if tool_call.name == "compile_final_score":
                    final_result = json.loads(tool_output)

                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": tool_call.id,
                    "content": tool_output
                })

            # Feed all tool results back to Claude in a single user turn
            messages.append({"role": "user", "content": tool_results})

            # Stop the loop if the model's stop reason signals it's finished
            if response.stop_reason == "end_turn":
                break

        return final_result if final_result else {"error": "Agent did not produce a final score"}
      
⚠️ Why we loop on stop_reason
Claude returns stop_reason: "tool_use" when it wants to call more tools and "end_turn" when it's finished. Always check this — if you break the loop too early, you'll get partial results. If you never break it, you could loop forever on a bug.

Step 4: Process Sample Leads and Generate Scores

Now let's wire everything together and run it against a couple of realistic leads. I'm using examples that reflect the kinds of businesses we work with here in Southwest Florida — a real estate brokerage and a restaurant group.

main.py
import json
from agent import LeadScoringAgent


def print_score_report(lead_name: str, result: dict) -> None:
    """Pretty-print the scoring result to the console."""
    print(f"\n{'='*60}")
    print(f"LEAD SCORE REPORT: {lead_name}")
    print(f"{'='*60}")
    print(f"Final Score:  {result['final_score']} / 100")
    print(f"Tier:         {result['tier']}")
    print(f"\nBreakdown:")
    for dimension, score in result['breakdown'].items():
        print(f"  {dimension.replace('_', ' ').title():<25} {score}")
    print(f"\nSummary:      {result['summary']}")
    print(f"Next Action:  {result['recommended_action']}")
    print(f"{'='*60}\n")


if __name__ == "__main__":

    agent = LeadScoringAgent()

    # Sample leads — swap these out for real form submissions or CRM records
    sample_leads = [
        {
            "name": "Maria Gonzalez",
            "title": "VP of Operations",
            "company": "Gulf Coast Realty Group",
            "industry": "Real Estate",
            "employees": "85",
            "budget": "$8,000–$12,000 per month",
            "timeline": "Want to go live before season starts — 6 weeks",
            "notes": "Interested in AI listing automation and lead routing. Currently using manual spreadsheets."
        },
        {
            "name": "Intern / Admin",
            "title": "Office Admin",
            "company": "Single-location sandwich shop",
            "industry": "Food & Beverage",
            "employees": "4",
            "budget": "Maybe a few hundred dollars",
            "timeline": "No rush, just curious what AI can do",
            "notes": "Found us on Google, not sure what they want yet."
        }
    ]

    for lead in sample_leads:
        result = agent.score_lead(lead)
        print_score_report(lead["name"], result)
      

Here's what the output actually looks like when you run this:

sample_output.txt
============================================================
LEAD SCORE REPORT: Maria Gonzalez
============================================================
Final Score:  87 / 100
Tier:         Hot

Breakdown:
  Budget                    23
  Company Size              16
  Decision Authority        18
  Timeline                  19
  Industry                  11

Summary:      Maria is a VP-level decision maker at an 85-person real estate firm
              with a clear budget and a hard deadline — all strong buying signals.
              Her use case maps directly to our listing automation offering.
Next Action:  Book discovery call within 24 hours. Prepare real estate AI demo.
============================================================

============================================================
LEAD SCORE REPORT: Intern / Admin
============================================================
Final Score:  14 / 100
Tier:         Disqualify

Breakdown:
  Budget                    3
  Company Size              2
  Decision Authority        2
  Timeline                  2
  Industry                  5

Summary:      This is a 4-person shop with no budget, no timeline, and no
              decision-making authority on the contact. Not a fit at this stage.
Next Action:  Add to newsletter list only. Do not assign to a sales rep.
============================================================
      

How It Works: Claude's Reasoning and Tool Selection

When Claude receives the lead information, it doesn't score blindly — it reads the rubric in the system prompt and decides which tools to call and in what sequence. It usually works dimension by dimension before calling compile_final_score, which forces it to show its work rather than just producing a number.

The tool-use architecture matters here because it makes the reasoning auditable. Every score has a reasoning field that tells you exactly why Claude landed where it did. If a sales manager disagrees with a score, you can look at the breakdown and adjust the rubric — not guess at what the model was thinking.

This is also why the agentic loop beats a single-prompt approach. A single prompt asking "score this lead from 1–100" gives you a number with no paper trail. The tool loop gives you structured data you can store, query, and improve over time.

Common Errors and Fixes

Error 1: anthropic.AuthenticationError

# Error message:
# anthropic.AuthenticationError: Error code: 401 - Authentication error

# Fix: Make sure your API key is actually in your environment
import os
print(os.environ.get("ANTHROPIC_API_KEY"))  # Should print your key, not None

# Set it before running:
# export ANTHROPIC_API_KEY="sk-ant-..."  (Mac/Linux)
# $env:ANTHROPIC_API_KEY="sk-ant-..."   (Windows PowerShell)
      

Error 2: Tool input validation error — missing required field

# Error message:
# anthropic.BadRequestError: tools.5.custom.input_schema: 
# 'reasoning' is a required property

# This happens when your tool schema lists a field as required 
# but Claude occasionally omits it. Fix: add default handling in your tool fn.

def score_budget_fit(stated_budget: str, reasoning: str = "Not provided", score: int = 0) -> dict:
    # Default values prevent crashes if Claude skips an optional reasoning field
    return {"dimension":