---
title: "News Sentiment Analysis in Python: VADER vs FinBERT (2026)"
description: "News sentiment analysis in Python: VADER vs FinBERT vs TextBlob on 200 real articles — agreement rates, title-vs-body flip, and a streaming loop."
source: https://apitube.io/blog/post/news-sentiment-analysis-python-vader-finbert-textblob
---

# News Sentiment Analysis in Python: VADER vs FinBERT vs TextBlob (2026)

Three libraries. One corpus. Honest numbers. Every tutorial I found scored headlines with a single library and called it done — then I tried to use that pipeline on real news body text and 40% of the verdicts flipped.

**News sentiment analysis in Python is the practice of assigning a polarity score (positive, neutral, or negative) to news text using NLP libraries such as VADER, TextBlob, or FinBERT, because downstream systems — trading signals, brand monitoring, risk dashboards — need a numeric signal rather than raw prose.** The right library depends on your corpus, latency budget, and whether you already have a news API that returns a built-in sentiment field.

This guide runs **VADER, TextBlob, and FinBERT on the same 200 real news articles** pulled from a live API, then compares the results against the API's built-in sentiment score. You'll finish with a working pipeline, three numbers that will change how you pick a library, and a production streaming loop.

**Disclosure**: I work on APITube. The code fetches articles from our `/everything` endpoint, but every library in this post (VADER, TextBlob, FinBERT) is API-agnostic — swap `requests.get` for NewsAPI, NewsCatcher, GNews, or a local CSV and the analysis code is identical.

## Key Takeaways

- On 200 tech-news articles, VADER and FinBERT agreed on the body-level label only 58% of the time. No single library is safe to ship alone.
- Title-only sentiment disagrees with body-level sentiment in ~40% of articles. Score bodies.
- FinBERT is ~3,000× slower than VADER on CPU. Use it only where domain jargon actually matters.
- If your news API returns a built-in sentiment that correlates with FinBERT at Pearson r ≥ 0.80 on your corpus, skip the custom model.
- Real-time news sentiment in Python is doable with a 5-minute polling loop, `published_at.start` windowing, and `id`-based dedupe — see Step 8.

## What You'll Build

By the end of this tutorial:

- A Python pipeline that pulls 200 English technology articles from the last 7 days
- Three sentiment scores per article (VADER, TextBlob, FinBERT) on both title and body
- A pairwise agreement table — how often do the libraries disagree?
- A title-vs-body flip rate — how often does body text reverse the verdict?
- A correlation between the API's built-in `sentiment.overall.score` and FinBERT — high r means you don't need a custom model
- A production streaming loop with dedupe, windowing, and a cost line per 1K articles

Total runtime on a MacBook Pro M2: ~8 minutes (FinBERT dominates).

## Prerequisites

```
python >= 3.10
pip install requests pandas vaderSentiment textblob transformers torch scipy
```

Get an APITube API key at apitube.io (free tier, 1K requests/month). Export it:

```bash
export APITUBE_KEY="sk_..."
```

The FinBERT model weights (ProsusAI/finbert on HuggingFace) are ~440 MB and download on first use.

## Step 1. Pull 200 Real News Articles

We want reproducible data: English, technology category, last 7 days. APITube's `/everything` endpoint takes all the filters we need.

```python
import os, requests, pandas as pd
from datetime import datetime, timedelta, timezone

APITUBE_KEY = os.environ["APITUBE_KEY"]
URL = "https://api.apitube.io/v1/news/everything"

def fetch_articles(n=200):
    end = datetime.now(timezone.utc)
    start = end - timedelta(days=7)
    rows, page = [], 1
    while len(rows) < n:
        r = requests.get(URL, headers={"X-API-Key": APITUBE_KEY}, params={
            "language.code": "en",
            "category.id": "medtop:13000000",
            "published_at.start": start.date().isoformat(),
            "published_at.end": end.date().isoformat(),
            "per_page": 50,
            "page": page,
        })
        r.raise_for_status()
        batch = r.json().get("results", [])
        if not batch:
            break
        for a in batch:
            body = a.get("body") or a.get("description") or ""
            if not body or len(body) < 140:
                continue
            rows.append({
                "id": a["id"],
                "title": a["title"],
                "body": body[:2000],
                "domain": a["source"]["domain"],
                "published_at": a["published_at"],
                "api_sentiment": a.get("sentiment", {}).get("overall", {}).get("score"),
            })
        page += 1
    return pd.DataFrame(rows[:n])

df = fetch_articles(200)
df.to_parquet("news_200.parquet")
print(df.shape, df.domain.nunique(), "unique domains")
```

Equivalent curl for a single page:

```bash
curl -H "X-API-Key: $APITUBE_KEY" \
  "https://api.apitube.io/v1/news/everything?language.code=en&category.id=medtop:13000000&per_page=50"
```

Cache the result to Parquet so you can rerun the scoring without burning API calls.

## Step 2. Score with VADER

VADER (Valence Aware Dictionary and sEntiment Reasoner) is lexicon and rule-based, built for short social text. It's fast: ~0.02 ms per sentence on CPU, no GPU needed.

```python
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
vader = SentimentIntensityAnalyzer()

def vader_score(text):
    return vader.polarity_scores(text)["compound"]  # range: -1..1

df["vader_title"] = df["title"].map(vader_score)
df["vader_body"]  = df["body"].map(vader_score)
```

The `compound` score is a normalized sum of valences in [-1, 1]. The conventional buckets: `>= 0.05` positive, `<= -0.05` negative, else neutral. VADER handles negation ("not good"), boosters ("very"), and all-caps, but it's weak on sarcasm and domain-specific language — it treats "bearish" as neutral.

## Step 3. Score with TextBlob

TextBlob uses the `pattern` polarity lexicon. Honestly: it's the weakest of the three for news. It has no negation handling and no intensity modifiers. I'm including it because it's popular in older tutorials — and so we can show where it fails.

```python
from textblob import TextBlob

df["tb_title"] = df["title"].map(lambda t: TextBlob(t).sentiment.polarity)
df["tb_body"]  = df["body"].map(lambda t: TextBlob(t).sentiment.polarity)
```

Output is polarity in [-1, 1], ignoring subjectivity. If you're shipping to production on TextBlob alone, stop. Every failure mode VADER has, TextBlob has worse.

## Step 4. Score with FinBERT

FinBERT is a BERT model fine-tuned on financial news (ProsusAI/finbert on HuggingFace). It outputs probabilities for three classes: positive / negative / neutral. We compress to a single signed score for comparison.

```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch, torch.nn.functional as F

tok = AutoTokenizer.from_pretrained("ProsusAI/finbert")
model = AutoModelForSequenceClassification.from_pretrained("ProsusAI/finbert")
model.eval()
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)

@torch.inference_mode()
def finbert_score(text):
    x = tok(text[:512], return_tensors="pt", truncation=True).to(device)
    probs = F.softmax(model(**x).logits, dim=-1)[0]
    # ProsusAI/finbert label order: positive, negative, neutral
    return float(probs[0] - probs[1])  # signed, range [-1, 1]

df["fb_title"] = df["title"].map(finbert_score)
df["fb_body"]  = df["body"].map(finbert_score)
```

Wall-clock numbers from my run on 200 articles (title + body each, so 400 inferences):

| Library | CPU (M2 Pro) | GPU (T4) |
|---|---|---|
| VADER | 0.04 s | — |
| TextBlob | 0.08 s | — |
| FinBERT | 118 s | 6.3 s |

Unlike VADER, which runs a deterministic lexicon lookup, FinBERT runs a full 110M-parameter transformer forward pass per article — which means on CPU it is ~3,000× slower. That gap dictates architecture choices the moment you scale past a few hundred articles. Truncate to 512 tokens (FinBERT's context limit) and batch in production (see Step 8).

## Step 5. Head-to-Head Agreement

Now the fun part: do they agree? I map each score to a three-way label (`pos` / `neg` / `neu`) at threshold 0.1 for VADER and TextBlob and 0.05 for FinBERT (tuned so the neutral-class rates are comparable), then count pairwise matches.

```python
def label(s, thr=0.1):
    if s >= thr: return "pos"
    if s <= -thr: return "neg"
    return "neu"

df["vader_label"]  = df["vader_body"].map(lambda s: label(s, 0.1))
df["tb_label"]     = df["tb_body"].map(lambda s: label(s, 0.1))
df["fb_label"]     = df["fb_body"].map(lambda s: label(s, 0.05))

pairs = [("vader_label", "tb_label"), ("vader_label", "fb_label"), ("tb_label", "fb_label")]
for a, b in pairs:
    agree = (df[a] == df[b]).mean()
    print(f"{a} vs {b}: {agree:.0%}")
```

Results on my 200-article tech corpus:

| Pair | Agreement |
|---|---|
| VADER ↔ TextBlob | 64% |
| VADER ↔ FinBERT | 58% |
| TextBlob ↔ FinBERT | 49% |

Unlike a synthetic benchmark (IMDB reviews, SST-2) where libraries are often reported at 85%+ agreement, this is real news body text — which means the ~40% of articles where any two libraries disagree are the ones you'd actually ship decisions on. Spot-check the disagreements before you trust any single library:

```python
diff = df[df.vader_label != df.fb_label].sample(5, random_state=0)
print(diff[["title", "vader_body", "fb_body"]])
```

A typical disagreement: a headline reads "Apple shares slump after disappointing guidance" — FinBERT labels it negative (correct), VADER labels it neutral because "slump" is not strongly weighted in its lexicon.

## Step 6. Title vs Body — Flip Rate

Every top-3 tutorial I read scores headlines. That is cheap but risky: headlines are often intentionally neutral or clickbait-positive while the body is negative.

```python
for lib in ["vader", "tb", "fb"]:
    thr = 0.05 if lib == "fb" else 0.1
    df[f"{lib}_title_lbl"] = df[f"{lib}_title"].map(lambda s: label(s, thr))
    df[f"{lib}_body_lbl"]  = df[f"{lib}_body"].map(lambda s: label(s, thr))
    flip = (df[f"{lib}_title_lbl"] != df[f"{lib}_body_lbl"]).mean()
    print(f"{lib} title→body flip rate: {flip:.0%}")
```

On my corpus:

| Library | Title → Body flip rate |
|---|---|
| VADER | 41% |
| TextBlob | 48% |
| FinBERT | 37% |

**In ~40% of articles the body tells a different story than the title.** If your downstream — trading, brand monitoring, risk — only sees the headline sentiment, you're wrong more than a third of the time. Score bodies.

## Step 7. Built-in vs Custom — Do You Need FinBERT?

APITube returns `sentiment.overall.score` on every article (a continuous score from an internal model). Here's a contrarian question: if it correlates highly with FinBERT, why run FinBERT at all?

```python
from scipy.stats import pearsonr
mask = df["api_sentiment"].notna()
r, p = pearsonr(df.loc[mask, "api_sentiment"], df.loc[mask, "fb_body"])
print(f"Pearson r = {r:.3f}, n = {mask.sum()}")
```

Result on my corpus: **Pearson r = 0.81** across 198 articles where the API returned a sentiment score. Decision framework:

- `r ≥ 0.80` → built-in is a reliable drop-in. Skip custom models unless you need domain-specific fine-tuning.
- `0.60 ≤ r < 0.80` → use built-in for coarse filtering; FinBERT for the top-N you care about.
- `r < 0.60` → the built-in is not aligned with your domain. Run a custom model or fine-tune one.

For most teams doing general news monitoring, built-in sentiment is good enough. The reason to run FinBERT is narrow: you need finance-specific domain language ("dovish", "guidance cut"), and even then, fine-tune on your own labels rather than using the off-the-shelf ProsusAI weights.

## Step 8. Production Streaming Loop

Tutorials stop after the one-shot fetch. Real systems need a streaming loop with windowed pulls, dedupe, and backoff. Here's the pattern:

```python
import time, json
from pathlib import Path

STATE = Path("state.json")

def load_cursor():
    if STATE.exists():
        return json.loads(STATE.read_text())["last_published_at"]
    return (datetime.now(timezone.utc) - timedelta(hours=1)).isoformat()

def save_cursor(ts):
    STATE.write_text(json.dumps({"last_published_at": ts}))

seen_ids = set()

def stream(poll_seconds=300):
    while True:
        cursor = load_cursor()
        r = requests.get(URL, headers={"X-API-Key": APITUBE_KEY}, params={
            "language.code": "en",
            "published_at.start": cursor,
            "per_page": 50,
        }, timeout=30)
        if r.status_code == 429:
            time.sleep(60); continue
        r.raise_for_status()
        batch = r.json().get("results", [])
        new = [a for a in batch if a["id"] not in seen_ids]
        if new:
            scores = [(a["id"], vader_score(a.get("body") or "")) for a in new]
            print(f"scored {len(scores)} new articles")
            seen_ids.update(a["id"] for a in new)
            latest = max(a["published_at"] for a in new)
            save_cursor(latest)
        time.sleep(poll_seconds)
```

Key production details:

- **Windowed cursor** — `published_at.start` is set to the latest `published_at` you've seen, not a fixed "last hour". Avoids missing late-arriving articles and avoids re-fetching old ones.
- **Dedupe by `id`** — APITube returns a stable article `id`. Keep a set (or a Redis SET in production).
- **429 backoff** — sleep 60 s, don't retry immediately.
- **VADER in the hot path** — FinBERT is too slow for sub-second latency; run it async on a queue for the articles VADER flags as strongly positive or negative.

**Cost line**: at $0.004 per request on APITube's Scale plan ($49/mo, 12,500 requests), 50 articles/request = $0.08 per 1K articles scored. Running FinBERT adds ~$0 on a small VM, but GPU inference on cloud (T4) is ~$0.06/hr → ~$0.05 per 1K articles at observed throughput. TextBlob and VADER are effectively free.

## Decision Framework: Which Library When?

No "it depends" wishy-washy. Use these thresholds:

| Scenario | Pick | Why |
|---|---|---|
| >1K articles/hour, need fast filter | **VADER** on body | 0.02 ms/article, handles negation |
| Finance-specific terms dominate | **FinBERT** + fine-tune | Off-the-shelf ProsusAI is for finance, not general news |
| Already use APITube and r > 0.80 with FinBERT on your corpus | **API built-in** | Skip custom model, save latency and $ |
| First prototype, <100 articles | VADER or TextBlob | Don't over-engineer |
| Production trading signal | Ensemble (VADER + FinBERT + API built-in), majority vote | Any single model is too noisy — our 49–64% agreement numbers are the proof |

Don't use TextBlob alone in 2026. Only include it if you need a sanity-check third opinion inside an ensemble.

## Frequently Asked Questions

### Which Python library is best for news sentiment analysis?

For general news body text, VADER is the best default: ~0.02 ms per sentence, handles negation and intensifiers, and agrees with FinBERT on 58% of body-level labels. Use FinBERT only when your corpus is financial or you plan to fine-tune. Avoid TextBlob as a standalone choice — it lacks negation handling and underperforms VADER on every test I ran.

### How accurate is VADER for news articles?

VADER is accurate for clear positive or negative headlines but weak on domain jargon and sarcasm. On a 200-article technology news corpus it agreed with FinBERT on 58% of body-level labels. Expect roughly 60–70% accuracy against human labels on general news; raise the decision threshold from 0.05 to 0.2 to reduce false positives at the cost of more neutrals.

### Is FinBERT better than VADER for financial news?

Yes, for finance-specific terms. FinBERT was fine-tuned on financial disclosures and correctly labels terms like "bearish", "guidance cut", or "dovish" that VADER misses. For general non-financial news, FinBERT's accuracy advantage narrows and its ~3,000× CPU latency penalty rarely justifies the switch. Fine-tune your own domain model before defaulting to ProsusAI weights.

### Can I do real-time sentiment analysis on news in Python?

Yes. Poll a news API with `published_at.start` set to the latest article timestamp you've seen, dedupe by article `id`, score with VADER in the hot path, and queue FinBERT asynchronously for articles flagged as strongly positive or negative. APITube's `/everything` endpoint supports 5-minute polling with sub-second response times; Step 8 above contains a working loop.

### How do you get news data for sentiment analysis in Python?

Use a news API that returns article `body`, not just headlines. Options: APITube (built-in sentiment, `body` field up to 2,000 chars), NewsAPI (headlines and description only on free tier), NewsCatcher, GNews. Avoid scraping — you hit paywalls and legal issues. Cache results to Parquet so reruns don't burn API calls against your monthly quota.

## Next Steps

You have enough to ship:

- Swap the category filter to your domain (`category.id=medtop:04000000` for business, `medtop:15000000` for sport, etc.)
- Run Pearson r between the built-in sentiment and FinBERT on **your** corpus before deciding if you need a custom model
- If you're doing trading signals, add domain-level aggregation — `df.groupby("domain").vader_body.mean()` — to spot publisher bias
- Wrap Step 8 in systemd or a Kubernetes CronJob for a real streaming service

**Try APITube free → [apitube.io](https://apitube.io)**. Free tier gives you 1K requests/month, which is enough to reproduce this entire tutorial.

## Resources

- **APITube** — [apitube.io](https://apitube.io) — try it free, sentiment and entities included on every article
- **Documentation** — [docs.apitube.io](https://docs.apitube.io) — endpoints, parameters, response structure, integrations
- **Pricing** — [apitube.io/pricing](https://apitube.io/pricing) — all tiers
- **APITube blog** — [apitube.io/blog](https://apitube.io/blog) — more guides and comparisons

**Related guides:**
- [Best News API for Python 2026](https://apitube.io/blog/best-news-api-python-2026)
- [Best News API for Sentiment & NLP 2026](https://apitube.io/blog/best-news-api-sentiment-nlp-2026)
