We pointed CatchAll Company Monitor at six watchlists (top VCs, top AI companies, this year’s two YC batches, and two lists shaped like our customers’), asked for every relevant news event in one week, and checked each result. CatchAll found 2.9× more verified events than an LLM web-search baseline and won every watchlist on F1. The main weakness is checking event dates.

This is a quick first run, and the idea is simple.

Our January and April benchmarks asked the question every search API gets: find all X that happened this week. Company Monitor answers the question customers ask next: what happened this week to the companies I care about?

You describe each company once: its name, website, a sentence about what it does, any other names it goes by, and its key people. You group the companies into a watchlist and attach it to a CatchAll job. You get back only events where one of your companies is a main actor, and each match comes with an explanation.

So we took companies picked around a few simple ideas:

  • Top 100 VCs. Investors whose every deal is news to someone.
  • Top 100 AI companies. The busiest news category there is.
  • Y Combinator’s Summer and Spring 2026 batches. 423 early-stage companies, many with short, common-word names.
  • Two lists modelled on our customers. One is a group of companies our clients actually track. The other mixes companies with a specific event type, the way customers set up monitors in practice.

Then we tried to find useful news about each of them, and checked every result. CatchAll did better on every list, by margins that ranged from narrow to overwhelming.

TL;DR

  • 6 watchlists, 692 companies, one week of news (1–7 September 2026), compared against an LLM web-search baseline
  • CatchAll found 705 verified events vs 245 (2.9×), with higher precision (0.534 vs 0.443) and much higher recall (0.819 vs 0.285)
  • CatchAll won all 6 watchlists on F1: 0.647 vs 0.347 overall
  • Its weak spot is checking event dates. 86% of CatchAll’s errors were real news about the right company, published that week, about an event from before the week. Counted by publication date rather than event date, CatchAll’s precision would be 0.937 and its F1 at least 0.824
  • For a monitor that runs daily, that means a feed that is almost entirely relevant news about the right companies
  • This is our first Company Monitor benchmark, so we compared against web search. Next time we’ll benchmark against a direct competitor
Watchlist Companies What we asked for
Customer group 24 All news
Top 100 AI companies 100 All news
Top 100 VCs 100 All news
YC Summer 2026 230 All news
YC Spring 2026 193 All news
Construction (customer-style) 45 Find all new construction projects, facility expansions and contract awards

6 watchlists · 692 companies · window 1–7 September 2026

Two kinds of monitor.

  • All news asks “tell me everything that happened to my companies this week”. There’s no topic filter, which makes it a direct test of whether the monitor can tell your company from everything else with a similar name.
  • Event monitors ask for one kind of event across the list.

We used both, because customers use both.

The customer group is a set of 24 companies our clients care about:

  • fintechs and banks (Revolut, Monzo, Wise, Banco BTG Pactual)
  • retailers (Shein, Temu, Oatly)
  • shipping lines (Zim, Wan Hai)
  • critical-minerals producers (Albemarle, Lynas, MP Materials)
  • several privately held companies that get very little press

It deliberately includes hard names: Bank of Georgia, which collides with the US state; Bark, a dog-products company that is also a dog noise; and a Brazilian bank covered mostly in Portuguese.

The construction list is the other common customer pattern, a company list paired with one event type. It mixes 25 large construction firms with 20 shipping, logistics and utility companies whose investment news often runs in French- and Chinese-language press.

How CatchAll was configured:

  • Base mode, with the watchlist attached through connected_dataset_ids.
  • All-news runs used fetch_all_watchlist_news, plus an event-date validator for the week.
  • The construction run used ed_association_type: event_associated, which returns only events where the company is an actor rather than a passing mention.

Both providers got the same two limits: 2 hours per query, and at most 1,000 records.

Methodology

How we chose the companies

  • Top 100 VCs and Top 100 AI companies: compiled with web research. Every firm needed a verified website, and any we couldn’t verify were dropped rather than guessed.
  • YC Summer and Spring 2026: taken from the batch lists, with founders listed for every company.
  • The two customer-style lists: built from the kinds of companies, sectors and event types our clients monitor.

Every provider received exactly the same description of each company: name, domain, description, aliases and key people.

How validation works

Company monitoring adds two questions to ordinary event checking:

  • Is this really my company? It might be another company with a similar name.
  • Is my company the subject of this story, or only mentioned in it?

So we score an (event, company) pair, not just an event.

The judge is Claude Sonnet 5 with live web search. It doesn’t take a provider’s text as evidence. It searches for the event itself and establishes:

  1. The event is real, and fits the query and the week. The date check uses the event’s own date, with two days’ tolerance either side.
  2. It’s the right company. The judge sees exactly the description the providers were given.
  3. The company is a main actor. The judge writes a one-sentence summary of the event, and we check whether the company appears in it. That framing gives far more consistent answers than asking “is this the main actor?” directly.

An event counts as correct when it passes the first check and at least one company it was linked to passes the other two. Every error gets a reason code, and no result was assumed correct for either provider.

Deduplication runs within each provider, so one event reported five times is judged once. It then runs across both providers, so we can see when both found the same event.

Recall is observable recall, as in our previous posts: measured against every verified event either provider found, not against everything that happened. True recall is lower for everyone, and that applies equally to both.

Why we compared against web search

This is our first attempt at benchmarking Company Monitor. The baseline is the obvious alternative: a strong LLM with web search, given the same company descriptions and asked for every event in the week.

We used OpenAI’s gpt-5-search-api, running one search per company, with follow-up rounds asking for anything it missed. A single prompt naming the whole list only covered a fraction of the companies. Asking about each company separately is the stronger baseline, and it’s the one we report.

It’s a strong baseline, but it isn’t a monitoring product. It has no scheduling, no deduplication between runs, and no explanation of why each company matched. It also needed 692 separate searches to cover these lists. Next time, we’ll benchmark against a real competitor.

Results

The short version

  • CatchAll finds far more of what happened to your companies. It found 2.9× as many verified events, at higher precision, and won on every watchlist.
  • Its weak spot is dates. Most of its mistakes were real, relevant news about the right company, describing something that happened before the week we asked about.
  • For a monitor you check every day, that weak spot matters much less. If you read your feed daily, nearly everything in it is relevant news about your companies.

How to read the numbers:

  • Precision: of everything a provider returned, the share that was correct.
  • Recall: of all verified events either provider found, the share this provider found.
  • F1: a single score balancing the two.
  • $/TP: what one verified, correct event cost.

Metrics are decimals, so 1.000 is perfect. Totals are weighted, so a watchlist with 500 events counts more than one with 5.

Provider F1 Precision Recall Verified events Watchlists won $/TP
CatchAll 0.647 ⭐ 0.534 ⭐ 0.819 ⭐ 705 ⭐ 6/6 ⭐ $0.202
Web-search baseline 0.347 0.443 0.285 245 0/6 $0.146 ⭐

Weighted totals · Observable universe: 861 unique verified events across 6 watchlists · ⭐ = Best in category

Watchlist Verified events (both) CatchAll F1 Baseline F1 CatchAll events Baseline events
Top 100 AI companies 563 0.649 ⭐ 0.290 479 122
Customer group 115 0.707 ⭐ 0.259 106 21
Top 100 VCs 115 0.587 ⭐ 0.541 76 63
Construction (customer-style) 34 0.650 ⭐ 0.500 26 15
YC Summer 2026 28 0.480 ⭐ 0.467 12 21
YC Spring 2026 6 0.750 ⭐ 0.214 6 3

6 watchlists · 1–7 September 2026 · ⭐ = Best F1 per watchlist

‍The busiest lists show the biggest gaps.

  • For the top 100 AI companies, CatchAll found 479 verified events to the baseline’s 122.
  • For the customer group, it found 106 to 21, and its recall was 0.922.

When a lot happens to a list, a search that pulls broadly and then filters to the list finds far more than checking company by company.

The VC list is closer. VC deals are announced with the investor’s name in the headline, so a targeted search per firm does well there.

The YC lists are thin. Most seed-stage companies don’t make the news in any given week. CatchAll still won both, but on very small numbers: 28 and 6 verified events between the two providers.

The construction list is the event-monitor case: one event type across a mixed, multilingual list. CatchAll found 26 projects and contract awards to the baseline’s 15, with recall of 0.765 against 0.441.

CatchAll’s weak spot: checking event dates

Every wrong result gets a reason. For CatchAll, one reason dominates:

Why results were wrong CatchAll Baseline
Event happened outside 1–7 September with ±2 days as allowable errors 531 289
Doesn’t match the query 39 9
Couldn’t identify the event 20 3
Event could not be verified as real 10 2
Other 14 5

False positives across the 6 watchlists

86% of CatchAll’s errors are out-of-window events. That’s 531 of 614. They are real news about the right company, published during the week, about something that happened earlier: a funding round from August reported again, or an acquisition closed the week before. Our benchmark is strict on purpose. If the event itself didn’t happen between 1 and 7 September, it counts as wrong.

CatchAll clearly has a problem validating event dates. We gave it an explicit event-date validator for this run, and it still let these through. Fixing that is our top priority.

It also shows how much quality is waiting to be unlocked. Here are the same results again, with out-of-window events no longer counted as errors:

Provider Precision F1 (at least)
CatchAll 0.534 → 0.937 0.647 → 0.824
Web-search baseline 0.443 → 0.966 0.347 → 0.478

The F1 figures are a conservative lower bound: we assume none of the recovered events overlap between the two providers.

Both providers improve, because the baseline’s errors are also mostly about timing. CatchAll’s lead holds: 0.824 against 0.478. On the two YC lists, the baseline would come out ahead under this looser rule, and that’s worth knowing: early-stage companies are where per-company search does best.

What this means if you run a monitor every day. A daily monitor shows you news published about your companies since yesterday. An article this week about a deal that closed last month is still news to you, and that’s exactly the kind of result our strict scoring counts as wrong. Measured that way, CatchAll’s feed is 94% relevant news about the right companies.

Other things worth knowing

  • The first ten results. Many people read a feed from the top. Across the six watchlists, 60% of CatchAll’s first ten results were correct, against 38% for the baseline. Position here is simply the order each provider returned results.
  • Cost. CatchAll cost $0.202 per verified event and the baseline $0.146. The baseline is cheaper per event because an LLM call is cheap, but it found only 28% of the verified events, against 82% for CatchAll.
  • Company coverage. The baseline returned at least one verified event for more of the listed companies: 235 against 159 out of 692. It searches every company individually, so it reaches more of the list. CatchAll finds more events, concentrated on the companies with the most going on.

Methodology notes

Recall is relative, not absolute. It’s measured against what the two providers found between them. True recall is lower for both and can’t be measured.

The baseline reads today’s web. CatchAll searched pages published during 1–7 September. The baseline’s web search has no date filter: it searched the web as it stood when we ran it, weeks later, and was asked to report only events from that week. That gives it material written after the fact, which an unfiltered search will always have.

One watchlist hit the result cap. CatchAll returned the maximum 1,000 records for the top 100 AI companies, so its recall there is understated.

Construction ran earlier. The construction list ran six days before the others, before we extended CatchAll’s search to include the last day of the window. Its search covered 1–6 September, which slightly understates CatchAll’s recall on that list.

Small lists, small numbers. The YC lists produced 28 and 6 verified events in total, so read their per-list numbers as directional.

CatchAll ran with its actor filter on for the construction list. That’s a real product feature, but a provider without it can’t match the precision gain it gives.

What’s next

  • A real competitor. This was a benchmark against web search. Next time we’ll run a monitoring product head to head.
  • Fixing event-date validation. This is the single biggest quality gain available.
  • Longer windows for early-stage lists. A week is too short to say much about seed-stage companies.

The full data is on Google Drive: Companies List & Queries. It contains:

  • the exact watchlists and queries
  • per-watchlist workbooks
  • every judged event, with the judge’s reasoning, the verified facts and citations

We’ll publish the next round whether the numbers go up or down.

Company Monitor is available in CatchAll today, and the documentation covers watchlists, datasets and connected jobs. Start with 2,000 free credits at platform.newscatcherapi.com. Questions: support@newscatcherapi.com.

Evaluation window: 1–7 September 2026 · Runs executed 24–30 September 2026 · Previous benchmarks: January 2026, April 2026

‍