This is the third post in our quarterly benchmark series. In January we showed CatchAll finding roughly 5× more relevant events than the closest competitor. In April its F1 rose from 0.527 to 0.705, winning 27 of 32 queries.

This quarter CatchAll's F1 is 0.605. That is lower than April, so we'll explain why before showing any results: we changed how results are verified, and the new method is much stricter. Here is what changed.

What changed since April

1. Validation now checks every result against the live web.

  • Until April: a fine-tuned LLM judge (92% agreement with our manual tagging) read the article text we extracted from each provider's URLs. When we couldn't retrieve a competitor's page, we counted the result as correct: 6–17% of competitor results in April.
  • Now: every result is judged by Claude Sonnet 5 with live web search. The judge doesn't trust the provider's text. It searches for the event itself and establishes three things:
    • Is the event real?
    • Did it happen inside the query's date range?
    • Does it match what was asked?
    If the open web can't settle it, a second pass looks at the sources the provider cited.

Real events now fail if they fall outside the date range, and results the provider made up or misdescribed get caught. Every provider's precision is lower as a result, and every number in this post is more trustworthy for it. April's figures and this quarter's aren't directly comparable. Our own numbers moved for the same reason everyone else's did.

2. OpenAI's deep-research model is gone. OpenAI shut down o3-deep-research on 23 July 2026. We now benchmark gpt-5-search-api, OpenAI's current search model, as the closest replacement. It's a fast search call rather than a long research agent, so it isn't the same product we tested before.

3. Every provider now gets the same two limits.

  • 2 hours per query. After that, we take whatever the provider has produced and score it.
  • 1,000 records per query. We send limit: 1000 to CatchAll, count: 1000 to Exa and match_limit: 1000 to Parallel AI. For the agent-based providers (Manus, OpenAI), we cap results at 1,000 when we collect them.

Nobody waits half a day for an answer in production, and an unlimited result set makes "more results" meaningless. The 2-hour limit is real: CatchAll hit it on 7 of 15 queries, and Exa on 6. Both were scored on partial results, which understates their recall.

4. A smaller round. Checking every event against the live web costs about $0.14 per event, on top of the language-model calls that group duplicates within each provider and then across providers. A round like April's 32 queries now costs four figures just to grade. So we ran a smaller set this quarter, and this post reports 15 queries:

  • 10 general event-discovery queries
  • 5 drawn from what our clients actually monitor

The full query list and raw data are available on request, as before.

5. The competitors.

Provider April Q3
Exa Websets Websets (unchanged)
Parallel AI Core generator Core generator (unchanged)
Manus 1.6 1.6 (unchanged)
OpenAI o3-deep-research gpt-5-search-api (o3 shut down)

We ran Base tier only this round. There are no Lite results.

TL;DR

  • 15 queries, 5 providers, one stricter judge that checks every result against the live web
  • CatchAll leads: F1 0.605, winning 10 of 15. Exa is next at 0.496 (4 wins), then Manus 0.362 (1 win), Parallel AI 0.264 and OpenAI 0.139
  • CatchAll found 836 verified events, 1.6× Exa's 525, at $0.188 each against Exa's $0.337
  • 96% of CatchAll's results are distinct events. For Exa it's 22%: about four in five Exa results repeat an event it already returned
  • New limits: 2 hours and 1,000 records per provider per query
  • For daily monitoring, where news published today counts even if the event happened last week, CatchAll's precision would be 0.744 and its F1 at least 0.688

Results

The short version

  • CatchAll finds the most of what actually happened. Its recall is 0.663; no one else reaches 0.42.
  • The precision-first providers are precise because they return little. Parallel AI, Manus and OpenAI all score above 0.82 precision, but find between 8% and 23% of the events.
  • Exa's record count overstates what it finds. It returned 3,941 records, which contained 855 distinct events.

How to read the numbers:

  • Precision: of everything a provider returned, the share that was verified correct.
  • Recall: of all the verified events any provider found, the share this provider found.
  • F1: a single score balancing the two.
  • $/TP: what one verified, correct event cost.

Totals are weighted, so a query with 250 events counts more than one with 6.

Provider F1 Precision Recall Verified events Query wins $/TP Unique results
CatchAll 0.605 ⭐ 0.557 0.663 ⭐ 836 ⭐ 10 / 15 ⭐ $0.188 95.4% ⭐
Exa Websets 0.496 0.614 0.416 525 4 / 15 $0.337 21.7%
Manus 1.6 0.362 0.825 0.232 292 1 / 15 $0.284 68.5%
Parallel AI Core 0.264 0.853 0.156 197 0 / 15 $0.349 89.2%
OpenAI gpt-5-search-api 0.139 0.864 ⭐ 0.075 95 0 / 15 $0.020 ⭐ 70.5%

Weighted totals · Observable universe: 1,261 unique verified events across 15 queries · ⭐ = Best in category

General event queries

Ten time-boxed questions of the kind any team asks a search API:

Query Window CatchAll Best competitor Winner
Completed M&A involving a US company 1–7 Sep 0.753 Exa 0.561 ✅ CatchAll
Material cyber incidents in SEC 8-K Item 1.05 filings 1–30 Sep 0.833 0.667 ✅ CatchAll
Facility or office closures in the US 1–7 Sep 0.656 Manus 0.368 ✅ CatchAll
Funding rounds raised by AI companies 1–7 Sep 0.643 Exa 0.585 ✅ CatchAll
Labour strikes worldwide 1–7 Sep 0.640 Manus 0.508 ✅ CatchAll
Regulatory fines issued to US companies 1–7 Sep 0.522 OpenAI 0.483 ✅ CatchAll
Layoffs of more than 50 at US companies 1–7 Sep 0.504 Parallel 0.400 ✅ CatchAll
Data-centre expansions or new builds 1–7 Sep 0.454 Exa 0.471 Exa
Workplace accidents in Taiwan 1–30 Sep 0.604 Exa 0.636 Exa
Series A rounds by US healthcare companies 1–7 Sep 0.500 Exa 0.636 Exa

‍CatchAll wins 7 of 10, with an F1 of 0.604 against Exa's 0.479. The widest margins are on high-volume questions:

  • Facility closures: 0.656 against a best competitor of 0.368.
  • M&A: 0.753 against 0.561.
  • Strikes: 0.640 against 0.508.

Two of the three losses are within a few hundredths: data centres by 0.017 and Taiwan by 0.033. The Taiwan query has been an Exa win in every round so far, so its result holds up across all three.

Client use-case queries

Five queries modelled on what our customers set up as monitors:

Query Window CatchAll Best competitor Winner
Profit warnings from LSE-listed companies 1–20 Sep 0.780 Manus 0.485 ✅ CatchAll
Anti-money-laundering enforcement against financial institutions 1–20 Sep 0.585 OpenAI 0.214 ✅ CatchAll
Disruptions or security incidents caused by a third-party vendor 1–7 Sep 0.495 Manus 0.333 ✅ CatchAll
Workplace fatalities at US company facilities 1–7 Sep 0.625 Manus 0.731 Manus
Fatal road-traffic collisions in Texas 1–7 Sep 0.657 Exa 0.897 Exa

‍CatchAll wins 3 of 5, with an F1 of 0.610 against Exa's 0.581.

The risk and compliance questions show the biggest gaps. On AML enforcement, Exa and Manus found nothing verifiable, and CatchAll scored 0.585. On third-party vendor incidents, CatchAll's 0.495 is well ahead of a best competitor at 0.333.

The two losses are small, local-incident questions, the pattern we've reported in every round:

  • Texas road fatalities: 90 events in total.
  • US workplace fatalities: 30 events in total.

With so few events, a precise search finds the core set cleanly, and a broader one picks up noise.

Exa's numbers include a lot of repeats

Exa returned the most records of any provider, 3,941. After grouping records that describe the same event, they contained 855 distinct events (21.7%).

The extreme case is CEO and CFO departures, a query outside this post's 15: Exa hit our 1,000-record cap with 37 distinct events. CatchAll's results are 95.4% distinct, because it groups articles into events before returning them.

This is partly by design. Exa returns web pages, one row per article, while CatchAll returns events. The effect for a user is the same either way, though: with Exa, you have to group the results into events yourself, and its record count overstates what it found by about 5×.

What changes for daily monitors

Our judge is strict about dates. If the event didn't happen inside the query's window, the result is wrong, even when the article was published that week. That's the right rule for a benchmark. It isn't how most people use a monitor.

Someone running a monitor every day reads what was published since yesterday. A story today about an acquisition closed last week is still news to them.

For CatchAll, 282 of its 666 false positives (42%) are exactly that: real events, correctly described, that happened just before the window. Counted as correct, the numbers become:

Provider Precision F1 (at least)
CatchAll 0.557 → 0.744 0.605 → 0.688
Exa Websets 0.614 → 0.812 0.496 → 0.533
Manus 1.6 0.825 → 0.890 0.362 → 0.299
Parallel AI Core 0.853 → 0.900 0.264 → 0.210
OpenAI 0.864 → 0.900 0.139 → 0.106

F1 is a conservative lower bound: we assume none of the recovered events overlap between providers. That gives the largest possible pool of events, and so the lowest recall for everyone.

CatchAll's lead grows under this reading: 0.688 against Exa's 0.533. For the three precision-first providers, F1 actually falls. Their precision barely moves, while the larger pool of events cuts their already-low recall.

Methodology notes

Recall is relative, not absolute. It's measured against what all five providers found between them, as in every round. True recall is lower for everyone and can't be measured.

This quarter isn't directly comparable with April. The judge, the time and volume limits, and the OpenAI model all changed. F1 is the most defensible number to compare across rounds, and even that moved because the method moved, not only because the products did.

Two providers were cut at the 2-hour limit. CatchAll was cut on 7 of 15 queries and Exa on 6. Both were scored on partial output, which understates their recall and slightly inflates everyone else's.

Forced true positives. When the judge can neither confirm nor refute an event, the result counts as correct rather than penalising the provider. This quarter that applied only to Manus: 38 results (13% of its true positives). Every other provider had none, down from 6–17% for competitors in April. Manus's precision is somewhat optimistic as a result.

Date tolerance. An event within two days of the window counts as in-window, because article and event dates often differ by a day or two.

The search APIs read today's web. Every provider searches the web as it is at run time. We ran these queries between 18 September and 1 October for windows in early-to-mid September, so material published after each window was available to all of them.

What's next

  • Event-date checking inside CatchAll. That's 42% of our false positives, and the largest precision gain available to us.
  • Enforcing narrow conditions. Tests outside this post's 15 showed CatchAll sometimes ignores a qualifier like "Fortune 500". That's the next most common source of error.
  • Back to quarterly. The next round will report the same metrics, whether the numbers go up or down.

Raw data and the full query list are available on request. Start with 2,000 free credits at platform.newscatcherapi.com. Questions: support@newscatcherapi.com.

Evaluation windows: 1–7 September 2026 (extended to 1–20 and 1–30 September for four queries) · Runs executed 18 September – 1 October 2026 · Previous benchmarks: January 2026, April 2026