#30 · Primary category: AI Content Readers & Aggregators
news-please
news-please - an integrated web crawler and information extractor for news that just works
Project last updated:04/14/26
GitHub Stars
2.5K
Forks
458
Contributors
47
License
Apache-2.0
Why we included this project
news-please saves you the tedium of scraping news sites yourself. Point it at a root URL and it follows internal links and RSS feeds to pull out the headline, body text, lead paragraph, authors, publication date, and language, then dumps the result into JSON, Elasticsearch, PostgreSQL, Redis, or wherever you want it. If you only need article extraction inside an existing pipeline, the library mode gives you that without spinning up a full crawl; the CLI is there for whole-site runs and can revisit articles to track revisions over time. For research teams assembling historical news corpora, the Common Crawl integration stands out, since it lets you filter that public archive by outlet and date instead of scraping each publisher one by one. The stack underneath is battle-tested Scrapy, Newspaper, and readability, so you get dependable extraction rather than experimental parsing logic.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
Folo
🧡 Folo is the AI RSS Reader
Vane
Vane is an AI-powered answering engine.
karakeep
A self-hostable bookmark-everything app (links, notes and images) with AI-based automatic tagging and full text search
daily
daily.dev is the personalized developer news feed and community. Get the best tech content from all over the web in your browser new tab or on mobile. Free and open source.
SurfSense
Open-source NotebookLM alternative. Research the open web with live data(Reddit, YT, IG, TikTok, Indeed, Google Search, Maps etc) through one platform, API or MCP server. Join our Discord: https://discord.gg/ejRNvftDp9