ProjectsSeptember 29, 2026

Web collection and sentiment pipeline

Built with
  • Python
Web collection and sentiment pipeline: illustrated cover

System architecture

Component and data flow diagramTopic configuration to Source collectors: Search terms. Source collectors to Visited URL set: Candidate URLs. Visited URL set to Article extraction: Unseen URLs. Article extraction to VADER: Extracted text. VADER to dataset.json: Append records.SEQUENTIAL PYTHON COLLECTION JOBSearch termsCandidate URLsUnseen URLsExtracted textAppend recordsTopic configurationEntities + keywordsSource collectorsWeb / news / RSSVisited URL setExact-URL dedupdataset.jsonCollected recordsVADERSentiment labelsArticle extractionnewspaper / HTML
Swipe horizontally to inspect the diagram.The pipeline attempts article extraction, then falls back to HTML paragraphs. Some sources provide headings only. Progress and deduplication are process-local.
Topic research spans multiple pages and feeds with inconsistent layouts. This script brings collected records into a common shape: source, title, URL, content and a sentiment label. A topic configuration expands categories into entities and search keywords. The pipeline calls collectors for news search, web search, Reddit headings and Google News RSS. Requests fetches HTML with a timeout; BeautifulSoup parses links and text, while feedparser handles RSS entries. For articles, newspaper extraction is attempted first. If it fails or produces insufficient text, paragraph extraction provides a fallback. Text is bounded before analysis. VADER assigns positive, negative or neutral labels from its compound score. A process-local visited set prevents repeated URL extraction, and the accumulated list is written to dataset.json. Sequential collection and a pause between keyword batches make the control flow easy to inspect. The cost is limited throughput and no durable progress if the run stops. An exact-URL set does not detect the same article under tracking links or syndicated copies. The Reddit collector uses headings rather than full article bodies, so not every sentiment label is based on the same amount of text. VADER is a lightweight baseline, not a validated measure of opinion for every topic or language. This is an extraction experiment with a JSON output, not a continuously operated crawler. HTML collectors can break when page structure changes or a source requires JavaScript. Several errors are swallowed, which makes a small output difficult to distinguish from failed collection. The next iteration would record per-source failures, preserve category and keyword provenance on every row, checkpoint progress and normalize URLs. I would prefer supported feeds or APIs where available and add per-source request policies before scheduling repeated runs.

Related projects

Let’s talk about the engineering

I’m open to software engineering roles across backend, platform, and data teams. Get in touch to discuss the architecture, trade-offs, or how this experience could help your team.
Get in touch