Summarize web articles: extract clean content, title, image, language, and generate text summaries
-
Updated
Jul 19, 2026 - Python
Summarize web articles: extract clean content, title, image, language, and generate text summaries
This repository is part of an NLP course for humanities and cultural studies. This course uses historical newspapers as a source and applies NLP methods to them. NLP tasks: Tokenization, Lemmatization, TF-IDF, Part-of-speech tagging, semantic search with transformers, article extraction and OCR post-correction with LLMs, NER and text classification
GNewsScraper is a TypeScript package that scrapes article data from Google News based on a keyword or phrase. It returns the results as an array of JSON objects, making it convenient to access and use the scraped information
强大易用的文章转 Markdown 工具,一键采集微信公众号、掘金、CSDN 等文章,自动下载图片、清理代码块、支持批量转换和 Agent Skill
A simple HTML-to-Markdown converter with article extraction, selector filtering, and batch conversion.
OpenClaw Skill:读取微信公众号文章、识别公众号并拉取文章列表
📋 WebMD is a Chrome extension that transforms web pages into Markdown documents with surgical precision.
A configurable pipeline for extracting and filtering articles from large corpora, tailored for the Delpher Kranten corpus, with support for features like keyword filtering and tf-idf-based relevance scoring.
A pipe-based news article scraping and metadata extraction library for Python
Mantis grabs exactly what you'd see on the page and turns it into structured JSON or clean Markdown. I built it for read-later and bookmarking tools, and for AI agents that need input without all the token bloat. It's Readability-style, has zero dependencies, and lives in a single file.
Chrome/Edge extension that estimates article word count, reading time, and lets you double-click any word to update the toolbar badge with your reading progress.
Zendesk articles extraction toolkit
HTML main-content extraction for Rust — ports of Mozilla Readability, Trafilatura, and htmldate.
Chrome extension that yoinks webpages into clean markdown. Supports article extraction, full-page capture, YouTube transcripts, and visual element picking.
AYLIEN is a news intelligence and text analysis platform providing REST APIs for article extraction, sentiment analysis, entity recognition, summarization, concept detection, and NLP-enriched news aggregation from over 80,000 public and licensed sources delivering 1.4 million articles daily.
India Times news extraction tool
US Magazine news extractor
CLI that extracts clean, readable article content from MHTML files — strips nav/ads/chrome, embeds images as data URIs, offline by default
Production web scraper with Playwright, bot-detection plugins, fingerprint rotation, and CAPTCHA solving. CLI + FastAPI.
Add a description, image, and links to the article-extraction topic page so that developers can more easily learn about it.
To associate your repository with the article-extraction topic, visit your repo's landing page and select "manage topics."