Trusted by 6,000+ developers

News API for RAG & Vector Pipelines

Retrieval is only as good as the corpus behind it. APITube feeds RAG pipelines full article bodies with structured metadata — entities, sentiment, categories, source authority, and story clusters — so you filter before you embed, cite what you retrieve, and keep the index fresh with real-time delivery.
Free Trial - No credit card required
Rated 5 stars by over 848 users
Realtime news articles from 177 countries
in 60 languages
  • CNN
  • Techcrunch
  • Vox
  • Apple
  • Microsoft
  • IBM
  • Bloomberg
  • Spotify

Benefits

A corpus built for retrieval, not a headlines feed

A news API for RAG supplies the documents a retrieval pipeline embeds, indexes, and serves to an LLM as citable context. APITube returns full article bodies from 300,000+ sources across 59 languages and 177+ countries — each with entities, per-entity sentiment, IPTC categories, source authority, and a story ID that groups duplicate coverage — exportable as JSONL or Parquet and kept fresh over SSE, WebSocket, or webhooks.

Get Free API Key
Learn more about API
1
Full article body in every response — ready to chunk and embed
2
Story clustering groups duplicate coverage before it pollutes your index
3
Entities across 9 types, linked to Wikipedia & Wikidata, with per-entity sentiment
4
Source authority (Open PageRank 0-10) and bias labels for corpus curation
5
JSONL & Parquet exports for batch ingestion
6
SSE, WebSocket & webhooks keep the index fresh without polling
7
Historical archive queryable by date range for backfills
8
Full enriched schema on the free plan

Discover the possibilities

Our solutions are designed to help you achieve your goals. Whether you are a small business or a large corporation, we have the right solution for you.

RAG

News-aware chat & Q&A

Answer questions about current events with retrieved, cited articles instead of a stale training set.

Agents

Research agents

Multi-step agents that retrieve, compare, and cite coverage across languages and sources.

Finance

Market & portfolio copilots

Ground financial Q&A in fresh, entity-tagged coverage with per-entity sentiment.

Risk

Compliance & risk retrieval

Retrieve negative-sentiment coverage about counterparties and suppliers at question time.

Media

Newsroom assistants

Summarize one story cluster instead of 40 duplicate articles, with sources attached.

Verification

Grounded fact-checking

Verify claims against the live index and return verdicts backed by ranked sources.

Features

Don’t waste time on complex features

Complex features made simple. Our API provides a simple way to access news articles from around the world. We provide a simple, consistent, and easy-to-use API to access news articles from thousands of sources.

News API

  • Export data in many formats
  • Industry Monitoring
  • Brand Monitoring
  • Market Intelligence
  • Risk Management
  • Competitive Intelligence
  • Media Monitoring
  • Sentiment Analysis
  • Trend Analysis
  • Story Grouping
  • Forecasting social trends
  • Multi-language Support
  • Audience Engagement
  • Geographical Analysis
  • Real-time Breaking News
  • Historical Data Access
  • Custom News Feeds
  • News Aggregation
  • Content Filtering
  • Over 50 integrations

Extract Additional Data

  • Industries
  • Locations
  • Persons
  • Organizations
  • Brands
  • Events
  • Disasters
  • Diseases
  • Links
  • Media
  • Images & Videos
  • Hashtags
  • Authors
  • Source
  • Duplicate Detection
  • Publisher Rank
  • Article Sentiment
  • Readability Score
  • Language Detection

Analysis

  • Sentiment Analysis
  • Analysis of public opinion
  • Categorization
  • Financial Analysis
  • Trend Analysis
  • Story Grouping
  • Content Summarization
  • Entity Recognition
  • Keyword Extraction
  • Topic Modeling
  • Event Detection
  • Named Entity Recognition
  • Text Classification
  • Controversy Detection
  • Trust Score Analysis
  • Engagement Metrics
  • Source Bias Detection
  • Quality Ranking
  • Spam Detection
  • Emotion Detection

Advanced Searching

  • Search by Location
  • Search by Date Range
  • Search by Source
  • Search by Category
  • Search by Industry
  • Search by Sentiment
  • Search by Story
  • Search by Publisher Rank
  • Search by Language
  • Search by Entity
  • Search by Keywords
  • Boolean Search
  • Proximity Search
  • Faceted Search
  • Range Queries
  • Search by Author
  • Search by Media Type
  • Search by Breaking News
  • Search by Read Time

More than 300,000+ sources

APITube is trusted by teams around the world to help them build and deliver amazing digital experiences faster than ever before.

Last updated
4s ago
Total sources
300.000k
Requests yesterday
20.888
Articles added yesterday
157.157k
Total articles
3.85b

Frequently asked questions

APITube is a news API built for retrieval-augmented generation. It returns full article bodies from 300,000+ sources across 59 languages, each enriched with entities, per-entity sentiment, IPTC categories, source authority, and a story endpoint that groups related coverage — exportable as JSONL or Parquet for embedding pipelines and kept fresh over SSE, WebSocket, or webhooks.
An LLM stops knowing the world at its training cutoff. A RAG pipeline closes that gap by retrieving current documents at question time — and news is the fastest-moving corpus there is. A news API gives the pipeline clean, structured, sourced articles to embed and cite, instead of scraped HTML with boilerplate and duplicates.
Yes. Every article from /v1/news/everything carries the full body along with the title and description, so you can chunk and embed real text rather than a headline. The same enriched schema is returned on every plan, including the free tier.
Three ways. Hold a stream open — Server-Sent Events or WebSocket — and embed articles as they arrive; register a webhook and let APITube POST new matches to your ingestion endpoint; or poll the REST API with date filters on a schedule. All channels accept the same filters, so the index only receives what your pipeline actually retrieves against.
Every APITube article carries a story ID that groups related coverage of one event into a single cluster. Ingest one representative article per story instead of 40 syndicated copies, and your retriever stops returning near-identical chunks for the same query.
Each APITube article carries entities across 9 types linked to Wikipedia and Wikidata, sentiment at the overall, title, body, and per-entity level from −1 to +1, IPTC Media Topics categories with relevance scores, the source with a political bias label and an Open PageRank authority score from 0 to 10, language, publish date, and a story ID. Store them as vector-database metadata and you can pre-filter retrieval by company, topic, sentiment, authority, or recency before similarity search runs.
The REST API exports JSON, JSONL, Parquet, CSV, TSV, XML, RSS, and XLSX. JSONL streams line-by-line into batch embedding jobs, and Parquet loads directly into the data warehouses and dataframe tools most ingestion pipelines already use.
Yes. APITube keeps a historical archive you can query by date range using published-at filters, so you can walk backwards through time windows and batch-ingest the depth your corpus needs before switching to real-time updates.
Not always. APITube runs a hosted MCP server at https://mcp.apitube.io/ that exposes news search as tools an agent can call directly — no vector database required. And the /v1/news/fact-check endpoint runs retrieval-augmented verification server-side: it checks a claim against the live index and returns a verdict with the coverage behind it.
Every article carries its canonical URL, source domain, source authority score, and publish date. Keep them in your chunk metadata and the model can attribute each statement to a linkable, ranked source — which is the difference between a grounded answer and a plausible one.

Related Solutions

Resources

Developers

We use cookies

By clicking "Accept", you agree to the storing of cookies on your device for functional and analytics.