# How to Detect Competitor Pricing Page Changes in Python Without False Alarms

To find out when a competitor changes their pricing page, fetch the page on a schedule, reduce it to its visible text, store that text, and diff it against the last snapshot. The diff is easy. The hard part is making it fire only when something real moved.

If you hash the raw HTML, you get an alert on every run. Asset fingerprints change, CSRF tokens rotate, a cookie banner gets a new line. None of it has anything to do with price. What follows is a small, dependency-free Python watcher that works on text instead of markup and drops the noise before it reaches you.

## Step one: extract only the visible text

A pricing change shows up in the text a person reads. Script bodies, inline styles and SVG paths are churn, so the parser skips them completely.

```python
import difflib
import re
import sqlite3
import urllib.request
from html.parser import HTMLParser

SKIP_TAGS = {"script", "style", "noscript", "svg", "template"}


class TextExtractor(HTMLParser):
 def __init__(self):
 super().__init__()
 self.skipping = []
 self.lines = []

 def handle_starttag(self, tag, attrs):
 if tag in SKIP_TAGS:
 self.skipping.append(tag)

 def handle_endtag(self, tag):
 if self.skipping and self.skipping[-1] == tag:
 self.skipping.pop()

 def handle_data(self, data):
 if self.skipping:
 return
 text = " ".join(data.split())
 if text:
 self.lines.append(text)
```

Collapsing whitespace inside each text node matters more than it looks. A reformatted template that only changes indentation would otherwise read as a change on every line.

## Step two: drop the lines that change on their own

After extraction you still have text that moves without anyone touching the prices: copyright years, "last updated" stamps, live visitor counters, consent banners. Put those in a list of patterns you own and filter them out before you store anything.

```python
VOLATILE = [
 re.compile(r"^©"),
 re.compile(r"^last updated", re.I),
 re.compile(r"\bviewing (this|now)\b", re.I),
 re.compile(r"\bcookies?\b", re.I),
]


def normalize(html):
 parser = TextExtractor()
 parser.feed(html)
 kept = [
 line for line in parser.lines
 if not any(p.search(line) for p in VOLATILE)
 ]
 return "\n".join(kept)
```

Grow this list from your own false positives. Each time the watcher fires on noise, the diff shows you exactly which line caused it, and that line becomes a pattern.

## Step three: treat a broken fetch as a failure, not a change

This failure mode is the one that wastes your time. The page comes back as a bot challenge, a login redirect, or an empty JavaScript shell. Diffed against yesterday's snapshot, it looks like the competitor deleted their entire pricing table.

The fix is a sanity check before the diff. A real pricing page contains something that looks like a price. If the fetched text has nothing like that, skip the run and keep the old snapshot.

```python
PRICE_MARKER = re.compile(
 r"[$€£]\s?\d|\bper (month|user|seat|year)\b|/mo\b|contact sales",
 re.I,
)


def looks_like_pricing(text):
 return bool(PRICE_MARKER.search(text))
```

Tune the marker for each competitor if you need to. The principle stays the same: a snapshot that fails the check never gets stored, so it can't poison the next comparison.

## Step four: store and diff in SQLite

One table and one row per URL. The diff only runs when the new text passed the sanity check and actually differs from what you stored.

```python
SCHEMA = """
CREATE TABLE IF NOT EXISTS snapshots (
 url TEXT PRIMARY KEY,
 body TEXT NOT NULL
)
"""


def fetch(url):
 req = urllib.request.Request(url, headers={"User-Agent": "pricing-watch"})
 with urllib.request.urlopen(req) as resp:
 return resp.read().decode(errors="replace")


def check(db, url):
 text = normalize(fetch(url))
 if not looks_like_pricing(text):
 return None # challenge page, empty shell or redirect: not a change

 row = db.execute(
 "SELECT body FROM snapshots WHERE url = ?", (url,)
 ).fetchone()

 db.execute(
 "INSERT INTO snapshots (url, body) VALUES (?, ?) "
 "ON CONFLICT(url) DO UPDATE SET body = excluded.body",
 (url, text),
 )
 db.commit()

 if row is None or row["body"] == text:
 return None

 return "\n".join(difflib.unified_diff(
 row["body"].splitlines(),
 text.splitlines(),
 fromfile="previous",
 tofile="current",
 lineterm="",
 ))


if __name__ == "__main__":
 db = sqlite3.connect("pricing.db")
 db.row_factory = sqlite3.Row
 db.execute(SCHEMA)
 diff = check(db, "https://competitor.example/pricing")
 if diff:
 print(diff)
```

In real use, pass a `timeout` to `urlopen` so a hanging server can't stall your scheduler. Run it from cron or any job runner you already have.

The output is a unified diff of human-readable lines. A removed line reading "Free plan" next to an added line reading "Free trial" tells you more than any hash ever will.

## Step five: confirm before you alert

If a competitor is A/B testing their pricing page, the watcher sees variant A on one run and variant B on the next, and fires both times. You can damp this by adding a `pending` column: when the text changes, write it to `pending` instead of `body`, and only promote it (and alert) if the next fetch returns the same text. That way you lose one cycle of latency and stop getting alerts that flap back and forth.

## The part a diff can't do

At this point you have a reliable answer to "did the visible text on this page change?" You don't have an answer to "does this matter to me?"

A renamed tier and a removed free plan both produce a few lines of diff. One is cosmetic. The other might be the opening your sales page should talk about this week. Telling them apart takes judgement about your own product, customers and positioning, and no regex has that. This is where it makes sense to hand the diff to an LLM along with a written description of your business and ask it how much this change matters to you specifically, not in general. The cheap deterministic layer decides that something changed. The expensive layer only runs on the survivors and decides whether you care.

## Limits of this approach

*   **JavaScript-rendered prices.** If a pricing page renders its numbers client-side, `urllib` gets the shell. The sanity check stops that from turning into a false alert, but it won't show you the prices. You would need a headless browser, which is a much heavier dependency.
    
*   **Personalized pricing.** Currency, region and logged-in state can all change what you see. You are watching the page your server gets, which may not be the page your customers get.
    
*   **Pages with nothing to diff.** A pricing page that only says "contact sales" has no numbers to track, and the change you care about happens in a sales call.
    
*   **Terms of service.** Check the site's terms and robots rules before you schedule fetches against it, and keep the frequency low. Pricing pages rarely need checking more than occasionally.
    
*   **Scale.** If you watch a single competitor, a calendar reminder to look at their page is honestly enough. This watcher earns its keep once there are more pages than you will remember to open.
    

Market Radar Kit is the self-hosted agent I built and run that wraps monthly competitor pricing-page tracking into eleven sources scored by Claude against your business context, and it needs a Docker host and your own Claude subscription: https://fulcrumenterprises.tech/go/market-radar-kit/?c=hashnode&v=8bd7cd
