News Sentiment Analysis in Python: VADER vs FinBERT (2026)

Erick Horn

Erick Horn

·

22 mins ler

News Sentiment Analysis in Python: VADER vs FinBERT (2026)

News Sentiment Analysis in Python: VADER vs FinBERT vs TextBlob (2026)

Three libraries. One corpus. Honest numbers. Every tutorial I found scored headlines with a single library and called it done — then I tried to use that pipeline on real news body text and 40% of the verdicts flipped.

News sentiment analysis in Python is the practice of assigning a polarity score (positive, neutral, or negative) to news text using NLP libraries such as VADER, TextBlob, or FinBERT, because downstream systems — trading signals, brand monitoring, risk dashboards — need a numeric signal rather than raw prose. The right library depends on your corpus, latency budget, and whether you already have a news API that returns a built-in sentiment field.

This guide runs VADER, TextBlob, and FinBERT on the same 200 real news articles pulled from a live API, then compares the results against the API's built-in sentiment score. You'll finish with a working pipeline, three numbers that will change how you pick a library, and a production streaming loop.

Disclosure: I work on APITube. The code fetches articles from our /everything endpoint, but every library in this post (VADER, TextBlob, FinBERT) is API-agnostic — swap requests.get for NewsAPI, NewsCatcher, GNews, or a local CSV and the analysis code is identical.

Key Takeaways

  • On 200 tech-news articles, VADER and FinBERT agreed on the body-level label only 58% of the time. No single library is safe to ship alone.
  • Title-only sentiment disagrees with body-level sentiment in ~40% of articles. Score bodies.
  • FinBERT is ~3,000× slower than VADER on CPU. Use it only where domain jargon actually matters.
  • If your news API returns a built-in sentiment that correlates with FinBERT at Pearson r ≥ 0.80 on your corpus, skip the custom model.
  • Real-time news sentiment in Python is doable with a 5-minute polling loop, published_at.start windowing, and id-based dedupe — see Step 8.

What You'll Build

By the end of this tutorial:

  • A Python pipeline that pulls 200 English technology articles from the last 7 days
  • Three sentiment scores per article (VADER, TextBlob, FinBERT) on both title and body
  • A pairwise agreement table — how often do the libraries disagree?
  • A title-vs-body flip rate — how often does body text reverse the verdict?
  • A correlation between the API's built-in sentiment.overall.score and FinBERT — high r means you don't need a custom model
  • A production streaming loop with dedupe, windowing, and a cost line per 1K articles

Total runtime on a MacBook Pro M2: ~8 minutes (FinBERT dominates).

Prerequisites

python >= 3.10
pip install requests pandas vaderSentiment textblob transformers torch scipy

Get an APITube API key at apitube.io (free tier, 1K requests/month). Export it:

export APITUBE_KEY="sk_..."

The FinBERT model weights (ProsusAI/finbert on HuggingFace) are ~440 MB and download on first use.

Step 1. Pull 200 Real News Articles

We want reproducible data: English, technology category, last 7 days. APITube's /everything endpoint takes all the filters we need.

import os, requests, pandas as pd
from datetime import datetime, timedelta, timezone

APITUBE_KEY = os.environ["APITUBE_KEY"]
URL = "https://api.apitube.io/v1/news/everything"

def fetch_articles(n=200):
    end = datetime.now(timezone.utc)
    start = end - timedelta(days=7)
    rows, page = [], 1
    while len(rows) < n:
        r = requests.get(URL, headers={"X-API-Key": APITUBE_KEY}, params={
            "language.code": "en",
            "category.id": "medtop:13000000",
            "published_at.start": start.date().isoformat(),
            "published_at.end": end.date().isoformat(),
            "per_page": 50,
            "page": page,
        })
        r.raise_for_status()
        batch = r.json().get("results", [])
        if not batch:
            break
        for a in batch:
            body = a.get("body") or a.get("description") or ""
            if not body or len(body) < 140:
                continue
            rows.append({
                "id": a["id"],
                "title": a["title"],
                "body": body[:2000],
                "domain": a["source"]["domain"],
                "published_at": a["published_at"],
                "api_sentiment": a.get("sentiment", {}).get("overall", {}).get("score"),
            })
        page += 1
    return pd.DataFrame(rows[:n])

df = fetch_articles(200)
df.to_parquet("news_200.parquet")
print(df.shape, df.domain.nunique(), "unique domains")

Equivalent curl for a single page:

curl -H "X-API-Key: $APITUBE_KEY" \
  "https://api.apitube.io/v1/news/everything?language.code=en&category.id=medtop:13000000&per_page=50"

Cache the result to Parquet so you can rerun the scoring without burning API calls.

Step 2. Score with VADER

VADER (Valence Aware Dictionary and sEntiment Reasoner) is lexicon and rule-based, built for short social text. It's fast: ~0.02 ms per sentence on CPU, no GPU needed.

from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
vader = SentimentIntensityAnalyzer()

def vader_score(text):
    return vader.polarity_scores(text)["compound"]  # range: -1..1

df["vader_title"] = df["title"].map(vader_score)
df["vader_body"]  = df["body"].map(vader_score)

The compound score is a normalized sum of valences in [-1, 1]. The conventional buckets: >= 0.05 positive, <= -0.05 negative, else neutral. VADER handles negation ("not good"), boosters ("very"), and all-caps, but it's weak on sarcasm and domain-specific language — it treats "bearish" as neutral.

Step 3. Score with TextBlob

TextBlob uses the pattern polarity lexicon. Honestly: it's the weakest of the three for news. It has no negation handling and no intensity modifiers. I'm including it because it's popular in older tutorials — and so we can show where it fails.

from textblob import TextBlob

df["tb_title"] = df["title"].map(lambda t: TextBlob(t).sentiment.polarity)
df["tb_body"]  = df["body"].map(lambda t: TextBlob(t).sentiment.polarity)

Output is polarity in [-1, 1], ignoring subjectivity. If you're shipping to production on TextBlob alone, stop. Every failure mode VADER has, TextBlob has worse.

Step 4. Score with FinBERT

FinBERT is a BERT model fine-tuned on financial news (ProsusAI/finbert on HuggingFace). It outputs probabilities for three classes: positive / negative / neutral. We compress to a single signed score for comparison.

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch, torch.nn.functional as F

tok = AutoTokenizer.from_pretrained("ProsusAI/finbert")
model = AutoModelForSequenceClassification.from_pretrained("ProsusAI/finbert")
model.eval()
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)

@torch.inference_mode()
def finbert_score(text):
    x = tok(text[:512], return_tensors="pt", truncation=True).to(device)
    probs = F.softmax(model(**x).logits, dim=-1)[0]
    # ProsusAI/finbert label order: positive, negative, neutral
    return float(probs[0] - probs[1])  # signed, range [-1, 1]

df["fb_title"] = df["title"].map(finbert_score)
df["fb_body"]  = df["body"].map(finbert_score)

Wall-clock numbers from my run on 200 articles (title + body each, so 400 inferences):

LibraryCPU (M2 Pro)GPU (T4)
VADER0.04 s—
TextBlob0.08 s—
FinBERT118 s6.3 s

Unlike VADER, which runs a deterministic lexicon lookup, FinBERT runs a full 110M-parameter transformer forward pass per article — which means on CPU it is ~3,000× slower. That gap dictates architecture choices the moment you scale past a few hundred articles. Truncate to 512 tokens (FinBERT's context limit) and batch in production (see Step 8).

Step 5. Head-to-Head Agreement

Now the fun part: do they agree? I map each score to a three-way label (pos / neg / neu) at threshold 0.1 for VADER and TextBlob and 0.05 for FinBERT (tuned so the neutral-class rates are comparable), then count pairwise matches.

def label(s, thr=0.1):
    if s >= thr: return "pos"
    if s <= -thr: return "neg"
    return "neu"

df["vader_label"]  = df["vader_body"].map(lambda s: label(s, 0.1))
df["tb_label"]     = df["tb_body"].map(lambda s: label(s, 0.1))
df["fb_label"]     = df["fb_body"].map(lambda s: label(s, 0.05))

pairs = [("vader_label", "tb_label"), ("vader_label", "fb_label"), ("tb_label", "fb_label")]
for a, b in pairs:
    agree = (df[a] == df[b]).mean()
    print(f"{a} vs {b}: {agree:.0%}")

Results on my 200-article tech corpus:

PairAgreement
VADER ↔ TextBlob64%
VADER ↔ FinBERT58%
TextBlob ↔ FinBERT49%

Unlike a synthetic benchmark (IMDB reviews, SST-2) where libraries are often reported at 85%+ agreement, this is real news body text — which means the ~40% of articles where any two libraries disagree are the ones you'd actually ship decisions on. Spot-check the disagreements before you trust any single library:

diff = df[df.vader_label != df.fb_label].sample(5, random_state=0)
print(diff[["title", "vader_body", "fb_body"]])

A typical disagreement: a headline reads "Apple shares slump after disappointing guidance" — FinBERT labels it negative (correct), VADER labels it neutral because "slump" is not strongly weighted in its lexicon.

Step 6. Title vs Body — Flip Rate

Every top-3 tutorial I read scores headlines. That is cheap but risky: headlines are often intentionally neutral or clickbait-positive while the body is negative.

for lib in ["vader", "tb", "fb"]:
    thr = 0.05 if lib == "fb" else 0.1
    df[f"{lib}_title_lbl"] = df[f"{lib}_title"].map(lambda s: label(s, thr))
    df[f"{lib}_body_lbl"]  = df[f"{lib}_body"].map(lambda s: label(s, thr))
    flip = (df[f"{lib}_title_lbl"] != df[f"{lib}_body_lbl"]).mean()
    print(f"{lib} title→body flip rate: {flip:.0%}")

On my corpus:

LibraryTitle → Body flip rate
VADER41%
TextBlob48%
FinBERT37%

In ~40% of articles the body tells a different story than the title. If your downstream — trading, brand monitoring, risk — only sees the headline sentiment, you're wrong more than a third of the time. Score bodies.

Step 7. Built-in vs Custom — Do You Need FinBERT?

APITube returns sentiment.overall.score on every article (a continuous score from an internal model). Here's a contrarian question: if it correlates highly with FinBERT, why run FinBERT at all?

from scipy.stats import pearsonr
mask = df["api_sentiment"].notna()
r, p = pearsonr(df.loc[mask, "api_sentiment"], df.loc[mask, "fb_body"])
print(f"Pearson r = {r:.3f}, n = {mask.sum()}")

Result on my corpus: Pearson r = 0.81 across 198 articles where the API returned a sentiment score. Decision framework:

  • r ≥ 0.80 → built-in is a reliable drop-in. Skip custom models unless you need domain-specific fine-tuning.
  • 0.60 ≤ r < 0.80 → use built-in for coarse filtering; FinBERT for the top-N you care about.
  • r < 0.60 → the built-in is not aligned with your domain. Run a custom model or fine-tune one.

For most teams doing general news monitoring, built-in sentiment is good enough. The reason to run FinBERT is narrow: you need finance-specific domain language ("dovish", "guidance cut"), and even then, fine-tune on your own labels rather than using the off-the-shelf ProsusAI weights.

Step 8. Production Streaming Loop

Tutorials stop after the one-shot fetch. Real systems need a streaming loop with windowed pulls, dedupe, and backoff. Here's the pattern:

import time, json
from pathlib import Path

STATE = Path("state.json")

def load_cursor():
    if STATE.exists():
        return json.loads(STATE.read_text())["last_published_at"]
    return (datetime.now(timezone.utc) - timedelta(hours=1)).isoformat()

def save_cursor(ts):
    STATE.write_text(json.dumps({"last_published_at": ts}))

seen_ids = set()

def stream(poll_seconds=300):
    while True:
        cursor = load_cursor()
        r = requests.get(URL, headers={"X-API-Key": APITUBE_KEY}, params={
            "language.code": "en",
            "published_at.start": cursor,
            "per_page": 50,
        }, timeout=30)
        if r.status_code == 429:
            time.sleep(60); continue
        r.raise_for_status()
        batch = r.json().get("results", [])
        new = [a for a in batch if a["id"] not in seen_ids]
        if new:
            scores = [(a["id"], vader_score(a.get("body") or "")) for a in new]
            print(f"scored {len(scores)} new articles")
            seen_ids.update(a["id"] for a in new)
            latest = max(a["published_at"] for a in new)
            save_cursor(latest)
        time.sleep(poll_seconds)

Key production details:

  • Windowed cursor — published_at.start is set to the latest published_at you've seen, not a fixed "last hour". Avoids missing late-arriving articles and avoids re-fetching old ones.
  • Dedupe by id — APITube returns a stable article id. Keep a set (or a Redis SET in production).
  • 429 backoff — sleep 60 s, don't retry immediately.
  • VADER in the hot path — FinBERT is too slow for sub-second latency; run it async on a queue for the articles VADER flags as strongly positive or negative.

Cost line: at $0.004 per request on APITube's Scale plan ($49/mo, 12,500 requests), 50 articles/request = $0.08 per 1K articles scored. Running FinBERT adds ~$0 on a small VM, but GPU inference on cloud (T4) is ~$0.06/hr → ~$0.05 per 1K articles at observed throughput. TextBlob and VADER are effectively free.

Decision Framework: Which Library When?

No "it depends" wishy-washy. Use these thresholds:

ScenarioPickWhy
>1K articles/hour, need fast filterVADER on body0.02 ms/article, handles negation
Finance-specific terms dominateFinBERT + fine-tuneOff-the-shelf ProsusAI is for finance, not general news
Already use APITube and r > 0.80 with FinBERT on your corpusAPI built-inSkip custom model, save latency and $
First prototype, <100 articlesVADER or TextBlobDon't over-engineer
Production trading signalEnsemble (VADER + FinBERT + API built-in), majority voteAny single model is too noisy — our 49–64% agreement numbers are the proof

Don't use TextBlob alone in 2026. Only include it if you need a sanity-check third opinion inside an ensemble.

Frequently Asked Questions

Which Python library is best for news sentiment analysis?

For general news body text, VADER is the best default: ~0.02 ms per sentence, handles negation and intensifiers, and agrees with FinBERT on 58% of body-level labels. Use FinBERT only when your corpus is financial or you plan to fine-tune. Avoid TextBlob as a standalone choice — it lacks negation handling and underperforms VADER on every test I ran.

How accurate is VADER for news articles?

VADER is accurate for clear positive or negative headlines but weak on domain jargon and sarcasm. On a 200-article technology news corpus it agreed with FinBERT on 58% of body-level labels. Expect roughly 60–70% accuracy against human labels on general news; raise the decision threshold from 0.05 to 0.2 to reduce false positives at the cost of more neutrals.

Is FinBERT better than VADER for financial news?

Yes, for finance-specific terms. FinBERT was fine-tuned on financial disclosures and correctly labels terms like "bearish", "guidance cut", or "dovish" that VADER misses. For general non-financial news, FinBERT's accuracy advantage narrows and its ~3,000× CPU latency penalty rarely justifies the switch. Fine-tune your own domain model before defaulting to ProsusAI weights.

Can I do real-time sentiment analysis on news in Python?

Yes. Poll a news API with published_at.start set to the latest article timestamp you've seen, dedupe by article id, score with VADER in the hot path, and queue FinBERT asynchronously for articles flagged as strongly positive or negative. APITube's /everything endpoint supports 5-minute polling with sub-second response times; Step 8 above contains a working loop.

How do you get news data for sentiment analysis in Python?

Use a news API that returns article body, not just headlines. Options: APITube (built-in sentiment, body field up to 2,000 chars), NewsAPI (headlines and description only on free tier), NewsCatcher, GNews. Avoid scraping — you hit paywalls and legal issues. Cache results to Parquet so reruns don't burn API calls against your monthly quota.

Next Steps

You have enough to ship:

  • Swap the category filter to your domain (category.id=medtop:04000000 for business, medtop:15000000 for sport, etc.)
  • Run Pearson r between the built-in sentiment and FinBERT on your corpus before deciding if you need a custom model
  • If you're doing trading signals, add domain-level aggregation — df.groupby("domain").vader_body.mean() — to spot publisher bias
  • Wrap Step 8 in systemd or a Kubernetes CronJob for a real streaming service

Try APITube free → apitube.io. Free tier gives you 1K requests/month, which is enough to reproduce this entire tutorial.

Resources

Related guides:

APITube - News API

Artigos relacionados

API de Notícias Sentimento 2026: 5 Fornecedores de NLP Comparados
Insights

API de Notícias Sentimento 2026: 5 Fornecedores de NLP Comparados

Análise de sentimento da API de Notícias em 2026: paridade de campo entre 5 fornecedores de NLP, diffs JSON reais, troca DIY vs. API e notas de confiabilidade de sentimento.

How to Detect Misinformation by Cross-Referencing Sources
Developer Guides

How to Detect Misinformation by Cross-Referencing Sources

Detect misinformation by cross-referencing sources: score corroboration breadth, source authority, and bias diversity over a news API. Code included.

Build Real-Time Market-Moving News Alerts (Python Tutorial)
Developer Guides

Build Real-Time Market-Moving News Alerts (Python Tutorial)

Build market-moving news alerts in Python: score, dedupe, and push the news that moves prices. Why sentiment is the wrong signal — with runnable code.

Earnings News Monitor: Build It in Python (2026)
Developer Guides

Earnings News Monitor: Build It in Python (2026)

Build an earnings news monitor in Python that fuses the earnings calendar with real-time company news over one watchlist — earnings-window framework + code.

Utilizamos cookies

Ao clicar em "aceitar", concorda com o armazenamento de cookies no seu dispositivo para fins funcionais e analíticos.