Python MIT

news-fetch

A Python Package which helps to scrape all news details from any news websites

S

santhoshse7en

Dernière activité 13 août 2026
santhoshse7en/news-fetch

230

étoiles

113

forks

0

issues ouvertes

extractsfelixgoogle-search-using-pythonkeywordlucasnewsnews-detailsnews-pleasenews-websitenewspaper3kpythonscraperscraper-engine

Ce README est souvent en anglais.

PyPI version Downloads Python versions License CI

news-fetch

Python news scraper & article extractor — extract title, text, authors, date, image, and publisher from any news URL. No API key. Confidence scores included.

Fetch news. Know why it worked.

pip install news-fetch
from newsfetch import fetch

article = fetch("https://www.thehindu.com/...")
print(article.title)
print(article.text)
print(article.authors)
print(article.published_at)
print(article.image)
print(article.confidence.overall)   # 0.0–1.0
print(article.content_source)       # e.g. json-ld.articleBody
news-fetch https://example.com/article
news-fetch https://example.com/article --json
news-fetch batch urls.txt -o articles.jsonl

Why news-fetch?

A lightweight alternative to newspaper3k / newspaper4k / trafilatura wrappers — with its own extraction engine, confidence scores, and bulk + proxy support.

Feature news-fetch
News article extraction (title, body, authors, date, image) ✅
Confidence scores + extraction provenance ✅
Bulk scraping (fetch_many / fetch_iter / CLI JSONL) ✅
Proxy + proxy rotation for thousands of URLs ✅
RSS / sitemap article discovery ✅
Async (pip install news-fetch[async]) ✅
Optional browser render (pip install news-fetch[browser]) ✅
Disk cache + robots.txt respect ✅
No API key / no account ✅
Small deps (lxml, requests, python-dateutil, cssselect) ✅

Install

pip install news-fetch
pip install news-fetch[async]     # httpx async fetch
pip install news-fetch[browser]   # Playwright fallback (then: playwright install chromium)

Requirements: Python 3.10+


Quick start

Single URL

from newsfetch import fetch

article = fetch(url)
print(article.title, article.text, article.confidence.overall)

From HTML (no network)

from newsfetch import extract

article = extract(html_bytes, url="https://example.com/story")

Bulk scraping + proxies

from newsfetch import fetch_many, fetch_iter

results = fetch_many(
    urls,
    max_workers=20,
    proxies=["http://user:pass@p1:8080", "http://user:pass@p2:8080"],
    request_delay=0.05,
)

for url, article in fetch_iter(urls, max_workers=16):
    if article:
        print(article.title)

Strict mode (production pipelines)

from newsfetch import fetch, LowConfidenceExtractionError

try:
    article = fetch(url, strict=True)
except LowConfidenceExtractionError as e:
    print(e.failed_fields, e.confidence.overall)

Discovery (RSS / sitemaps)

from newsfetch import discover

for item in discover("https://www.bbc.com", limit=10):
    print(item["url"], item.get("title"))

Async

from newsfetch import fetch_async, fetch_many_async

article = await fetch_async(url)
articles = await fetch_many_async(urls, max_concurrency=50, proxies=PROXIES)

CLI

news-fetch https://example.com/article
news-fetch get URL --json
news-fetch batch urls.txt -o out.jsonl --workers 20
news-fetch discover https://www.theguardian.com --limit 10

Cache / robots / browser

from newsfetch import fetch, Config, NewsFetcher

fetch(url, cache=True, respect_robots=True)
fetch(url, render=True)                      # needs news-fetch[browser]
fetch(url, browser_fallback=True)            # retry with Playwright if confidence is low

Custom strategy plugin

from newsfetch import NewsFetcher, CallableStrategy
from newsfetch.strategies.base import Candidate

def my_strategy(doc):
    return {"title": [Candidate("Custom", "plugin.custom", 0.99)]}

fetcher = NewsFetcher()
fetcher.register_strategy(CallableStrategy("custom", my_strategy))

Article fields

url · canonical_url · title · description · text · authors · published_at · modified_at · publisher · language · image · keywords · section · summary · word_count · reading_time_minutes · page_type · is_article · confidence · extraction · sources

article.to_dict()
article.to_json()

MIT License · Built for developers who need reliable Python news scraping without an API.

Projets similaires

newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:

Pythoncrawlercrawlingnews
Ccodelucas
15,2 k étoiles2,1 k

news-please - an integrated web crawler and information extractor for news that just works

Pythoncc-newsccnewscommoncrawl
Ffhamborg
2,5 k étoiles458

📰 Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.

Pythonarticlesarticles-datacrawler
AAndyTheFactory
1,1 k étoiles116