This is the third post in our quarterly benchmark series. In January we showed CatchAll finding roughly 5× more relevant events than the closest competitor. In April its F1 rose from 0.527 to 0.705, winning 27 of 32 queries.
This quarter CatchAll's F1 is 0.605. That is lower than April, so we'll explain why before showing any results: we changed how results are verified, and the new method is much stricter. Here is what changed.
What changed since April
1. Validation now checks every result against the live web.
- Until April: a fine-tuned LLM judge (92% agreement with our manual tagging) read the article text we extracted from each provider's URLs. When we couldn't retrieve a competitor's page, we counted the result as correct: 6–17% of competitor results in April.
- Now: every result is judged by Claude Sonnet 5 with live web search. The judge doesn't trust the provider's text. It searches for the event itself and establishes three things:
- Is the event real?
- Did it happen inside the query's date range?
- Does it match what was asked?
Real events now fail if they fall outside the date range, and results the provider made up or misdescribed get caught. Every provider's precision is lower as a result, and every number in this post is more trustworthy for it. April's figures and this quarter's aren't directly comparable. Our own numbers moved for the same reason everyone else's did.
2. OpenAI's deep-research model is gone. OpenAI shut down o3-deep-research on 23 July 2026. We now benchmark gpt-5-search-api, OpenAI's current search model, as the closest replacement. It's a fast search call rather than a long research agent, so it isn't the same product we tested before.
3. Every provider now gets the same two limits.
- 2 hours per query. After that, we take whatever the provider has produced and score it.
- 1,000 records per query. We send
limit: 1000to CatchAll,count: 1000to Exa andmatch_limit: 1000to Parallel AI. For the agent-based providers (Manus, OpenAI), we cap results at 1,000 when we collect them.
Nobody waits half a day for an answer in production, and an unlimited result set makes "more results" meaningless. The 2-hour limit is real: CatchAll hit it on 7 of 15 queries, and Exa on 6. Both were scored on partial results, which understates their recall.
4. A smaller round. Checking every event against the live web costs about $0.14 per event, on top of the language-model calls that group duplicates within each provider and then across providers. A round like April's 32 queries now costs four figures just to grade. So we ran a smaller set this quarter, and this post reports 15 queries:
- 10 general event-discovery queries
- 5 drawn from what our clients actually monitor
The full query list and raw data are available on request, as before.
5. The competitors.
We ran Base tier only this round. There are no Lite results.
TL;DR
- 15 queries, 5 providers, one stricter judge that checks every result against the live web
- CatchAll leads: F1 0.605, winning 10 of 15. Exa is next at 0.496 (4 wins), then Manus 0.362 (1 win), Parallel AI 0.264 and OpenAI 0.139
- CatchAll found 836 verified events, 1.6× Exa's 525, at $0.188 each against Exa's $0.337
- 96% of CatchAll's results are distinct events. For Exa it's 22%: about four in five Exa results repeat an event it already returned
- New limits: 2 hours and 1,000 records per provider per query
- For daily monitoring, where news published today counts even if the event happened last week, CatchAll's precision would be 0.744 and its F1 at least 0.688
Results
The short version
- CatchAll finds the most of what actually happened. Its recall is 0.663; no one else reaches 0.42.
- The precision-first providers are precise because they return little. Parallel AI, Manus and OpenAI all score above 0.82 precision, but find between 8% and 23% of the events.
- Exa's record count overstates what it finds. It returned 3,941 records, which contained 855 distinct events.
How to read the numbers:
- Precision: of everything a provider returned, the share that was verified correct.
- Recall: of all the verified events any provider found, the share this provider found.
- F1: a single score balancing the two.
- $/TP: what one verified, correct event cost.
Totals are weighted, so a query with 250 events counts more than one with 6.
General event queries
Ten time-boxed questions of the kind any team asks a search API:
CatchAll wins 7 of 10, with an F1 of 0.604 against Exa's 0.479. The widest margins are on high-volume questions:
- Facility closures: 0.656 against a best competitor of 0.368.
- M&A: 0.753 against 0.561.
- Strikes: 0.640 against 0.508.
Two of the three losses are within a few hundredths: data centres by 0.017 and Taiwan by 0.033. The Taiwan query has been an Exa win in every round so far, so its result holds up across all three.
Client use-case queries
Five queries modelled on what our customers set up as monitors:
CatchAll wins 3 of 5, with an F1 of 0.610 against Exa's 0.581.
The risk and compliance questions show the biggest gaps. On AML enforcement, Exa and Manus found nothing verifiable, and CatchAll scored 0.585. On third-party vendor incidents, CatchAll's 0.495 is well ahead of a best competitor at 0.333.
The two losses are small, local-incident questions, the pattern we've reported in every round:
- Texas road fatalities: 90 events in total.
- US workplace fatalities: 30 events in total.
With so few events, a precise search finds the core set cleanly, and a broader one picks up noise.
Exa's numbers include a lot of repeats
Exa returned the most records of any provider, 3,941. After grouping records that describe the same event, they contained 855 distinct events (21.7%).
The extreme case is CEO and CFO departures, a query outside this post's 15: Exa hit our 1,000-record cap with 37 distinct events. CatchAll's results are 95.4% distinct, because it groups articles into events before returning them.
This is partly by design. Exa returns web pages, one row per article, while CatchAll returns events. The effect for a user is the same either way, though: with Exa, you have to group the results into events yourself, and its record count overstates what it found by about 5×.
What changes for daily monitors
Our judge is strict about dates. If the event didn't happen inside the query's window, the result is wrong, even when the article was published that week. That's the right rule for a benchmark. It isn't how most people use a monitor.
Someone running a monitor every day reads what was published since yesterday. A story today about an acquisition closed last week is still news to them.
For CatchAll, 282 of its 666 false positives (42%) are exactly that: real events, correctly described, that happened just before the window. Counted as correct, the numbers become:
CatchAll's lead grows under this reading: 0.688 against Exa's 0.533. For the three precision-first providers, F1 actually falls. Their precision barely moves, while the larger pool of events cuts their already-low recall.
Methodology notes
Recall is relative, not absolute. It's measured against what all five providers found between them, as in every round. True recall is lower for everyone and can't be measured.
This quarter isn't directly comparable with April. The judge, the time and volume limits, and the OpenAI model all changed. F1 is the most defensible number to compare across rounds, and even that moved because the method moved, not only because the products did.
Two providers were cut at the 2-hour limit. CatchAll was cut on 7 of 15 queries and Exa on 6. Both were scored on partial output, which understates their recall and slightly inflates everyone else's.
Forced true positives. When the judge can neither confirm nor refute an event, the result counts as correct rather than penalising the provider. This quarter that applied only to Manus: 38 results (13% of its true positives). Every other provider had none, down from 6–17% for competitors in April. Manus's precision is somewhat optimistic as a result.
Date tolerance. An event within two days of the window counts as in-window, because article and event dates often differ by a day or two.
The search APIs read today's web. Every provider searches the web as it is at run time. We ran these queries between 18 September and 1 October for windows in early-to-mid September, so material published after each window was available to all of them.
What's next
- Event-date checking inside CatchAll. That's 42% of our false positives, and the largest precision gain available to us.
- Enforcing narrow conditions. Tests outside this post's 15 showed CatchAll sometimes ignores a qualifier like "Fortune 500". That's the next most common source of error.
- Back to quarterly. The next round will report the same metrics, whether the numbers go up or down.
Raw data and the full query list are available on request. Start with 2,000 free credits at platform.newscatcherapi.com. Questions: support@newscatcherapi.com.
Evaluation windows: 1–7 September 2026 (extended to 1–20 and 1–30 September for four queries) · Runs executed 18 September – 1 October 2026 · Previous benchmarks: January 2026, April 2026



























.png)



.png)







































