# Tuning Keyword Weights From Thumbs Up/Down Feedback in a Monitoring Agent

A keyword pre-filter is the cheap half of any two-pass monitoring pipeline. Items arrive from your sources, the filter matches terms you own, sums their weights, and decides which survivors are worth an LLM call. The weight table is the entire policy. And on the day you write it, every number in it is a guess.

The guess also goes stale. Your market renames things. A term that used to signal a real threat starts appearing in every third post, and now it is noise that costs you tokens. A term you never thought to weight heavily turns out to sit on top of everything you actually read.

The fix is a feedback loop: you vote on items you see, and a tuner turns those votes into proposed changes to the weight table. Here is how to build it, and the two places it quietly goes wrong.

## Step one: log which keywords fired, not just the score

This is the part that is painful to retrofit. If your filter stores only a final score on the item row, you have no way to attribute a vote back to the terms that caused the item to surface. You need the match record written at scoring time.

```sql
CREATE TABLE keyword (
 term TEXT PRIMARY KEY,
 dimension TEXT NOT NULL,
 weight REAL NOT NULL
);

CREATE TABLE keyword_hit (
 item_id INTEGER NOT NULL,
 term TEXT NOT NULL,
 dimension TEXT NOT NULL,
 PRIMARY KEY (item_id, term)
);

CREATE TABLE feedback (
 item_id INTEGER PRIMARY KEY,
 vote INTEGER NOT NULL, -- +1 or -1
 created_at TEXT NOT NULL
);
```

The filter writes a `keyword_hit` row per match as it scores:

```python
def score_item(conn, item_id, text, keywords):
 total = 0.0
 hits = []
 haystack = text.lower()
 for term, dimension, weight in keywords:
 if term in haystack:
 hits.append((item_id, term, dimension))
 total += weight
 conn.executemany(
 "INSERT OR IGNORE INTO keyword_hit VALUES (?, ?, ?)", hits
 )
 return total, hits
```

Writing hits costs you a few rows per item and buys you every tuning decision you will ever make. Without it, a thumbs-down is just a bit on an item you will never see again.

## Step two: turn votes into per-term precision

Now a vote on an item becomes evidence about each term that fired on it. The naive aggregation is a join and a group by:

```sql
SELECT kh.term,
 SUM(CASE WHEN f.vote > 0 THEN 1 ELSE 0 END) AS up,
 SUM(CASE WHEN f.vote < 0 THEN 1 ELSE 0 END) AS down
FROM keyword_hit kh
JOIN feedback f USING (item_id)
GROUP BY kh.term;
```

Raw counts are a bad basis for a weight change, because a term with one up vote and nothing else looks perfect. Smooth it, and refuse to propose anything for a term without enough votes behind it:

```python
MIN_SUPPORT = 8
ALPHA = 1.0
MAX_STEP = 0.25

def propose(term, current, up, down):
 support = up + down
 if support < MIN_SUPPORT:
 return None
 precision = (up + ALPHA) / (support + 2 * ALPHA)
 # precision 0.5 is neutral; scale the weight around it
 factor = 0.5 + precision
 target = current * factor
 low, high = current * (1 - MAX_STEP), current * (1 + MAX_STEP)
 return {"term": term, "from": current,
 "to": round(min(max(target, low), high), 3),
 "support": support, "precision": round(precision, 3)}
```

`MIN_SUPPORT` is your patience knob and `MAX_STEP` is your safety knob. The step clamp matters more than it looks: it means a bad week of voting cannot flatten a term to zero, and you keep getting evidence about it instead of blinding yourself to it.

## Credit assignment: one vote lands on every term that fired

An item rarely fires one term. It fires a cluster, and your vote lands on all of them equally. Vote down a long press release that happened to mention your competitor, and the competitor's name takes the same penalty as the boilerplate term that actually made it noise.

Two mitigations you can apply, both standard:

**Split the credit.** Divide each vote by the number of terms that fired on that item, so a crowded item carries less weight per term than a clean one.

```sql
WITH hit_count AS (
 SELECT item_id, COUNT(*) AS n FROM keyword_hit GROUP BY item_id
)
SELECT kh.term,
 SUM(CASE WHEN f.vote > 0 THEN 1.0 / hc.n ELSE 0 END) AS up,
 SUM(CASE WHEN f.vote < 0 THEN 1.0 / hc.n ELSE 0 END) AS down
FROM keyword_hit kh
JOIN feedback f USING (item_id)
JOIN hit_count hc USING (item_id)
GROUP BY kh.term;
```

**Compare against the term's own base rate.** A term that fires on nearly everything in the corpus carries almost no information, whatever its vote record. Compute how often it matches overall, and treat the high-frequency terms as candidates for a weight cut rather than a boost, regardless of which way the votes went.

## Propose, do not apply

The tempting design is a job that writes new weights straight into the table overnight. Do not do that. Three reasons, all of them boring and all of them real.

Your labels are skewed by your own UI. If a thumbs-down is the easy click on an alert and a thumbs-up requires opening the item, your negative evidence will outweigh your positive evidence for reasons that have nothing to do with relevance.

Auto-applying narrows the funnel in a direction you cannot see. Weights go down on terms attached to items you skimmed and up on terms attached to the topic you happen to be deep in this month. A few cycles of that and your filter is an echo of last month's attention, with no record of what it stopped showing you.

And a weight table is a document you need to be able to reason about. If it changes on its own, debugging a missed item turns into archaeology.

A proposal is cheap to review. Emit it as a diff with the evidence attached, and keep the apply step a human keystroke:

```python
for p in proposals:
 print(f"{p['term']:28} {p['from']:>6} -> {p['to']:<6} "
 f"support={p['support']} precision={p['precision']}")
```

Two useful guardrails on the reviewer's side: a cooldown so the same term cannot move on consecutive runs, and an audit row for every applied change, so you can answer "when did this term get cheap, and on whose vote" later.

## What this does not fix

Tuning the cheap pass is a narrow win, and it is worth being clear about the edges.

It cannot fix your source list. If the thing you needed never entered the pipeline, no weight change will surface it. Feedback tuning sharpens the ranking of what you already ingest, nothing more.

It cannot find vocabulary you never wrote down. A tuner that adjusts weights on existing terms will never propose a term that is absent from the table. New language has to come from you reading the raw firehose sometimes, or from the expensive pass telling you what it keeps seeing.

It needs you to actually vote. Support thresholds are doing real work here, and until a term crosses one it just sits at its initial guess. If nobody on the team will click, skip the whole mechanism and hand-edit the weights when they annoy you. That is a legitimate choice and a smaller amount of code.

It is also the wrong tool if your volume is low enough to read in full. A pre-filter earns its keep when the firehose is bigger than your attention. Below that line, the filter is just a place for bugs to hide.

I build and run Market Radar Kit, a self-hosted market-intelligence agent with 11 ingestion sources, a four-dimension keyword pass in front of Claude scoring, and an auto-tuner that proposes keyword-weight adjustments from your thumbs up/down feedback. The keyword pass uses no LLM; the Claude scoring layer runs through the Claude Code CLI on your own Claude subscription, which the kit requires: https://fulcrumenterprises.tech/go/market-radar-kit/?c=hashnode&v=8bd7cd
