Skip to main content
As part of the NewsCatcher processing pipeline, each article is enriched with NLP data before it is indexed: theme classification, sentiment scores, named entities, content tags, and vector embeddings. News API exposes these fields in the response when you set include_nlp_data to true.
include_nlp_data applies to the News API only. On the Local News API, the nlp object is returned by default on every subscription plan, and sending include_nlp_data returns 403 Param include_nlp_data is not allowed for plan <plan_id>.
NLP enrichment is available only for articles indexed from July 2023 onward. For earlier articles, the API returns "nlp": {}.To request NLP enrichment for historical articles, contact support@newscatcherapi.com.

How NLP processing works

Processing mode depends on the article’s language and determines which response fields are populated and which are null. Native processing applies to English and Arabic articles. NLP runs on the original text and results appear in the standard nlp.* fields. Translation-based processing applies to all other languages. The article is first translated to English, then NLP runs on that translation. Results appear in nlp.translation_* fields — the corresponding standard fields are explicitly null, not absent. To receive translation fields in the response, set include_translation_fields to true. This distinction matters when consuming NER or summary fields: a null value in nlp.ner_PER means the article was processed via translation, not that no entities exist — check nlp.translation_ner_PER instead.

Available features

See also