03 / Research Automation & Information Analysis

Automotive News Intelligence Workflow

Rule-based news research and source review

The problem

I wanted an easier way to follow relevant automotive and EV developments without getting lost in the volume of news.

Personal portfolio case · Stored RSS snapshot

Source & implementation

Why I built it

Following automotive news made it difficult to separate useful developments from noise. I explored whether collecting articles, removing duplicates and organising them by brand and topic would make the material easier to review.

I wanted to search by brand or topic and open the original articles when a result needed checking.

My contribution

I built a workflow around a stored RSS snapshot: collect records, remove duplicate URLs, apply a relevance filter and group the remaining records with keyword rules. The dashboard lets me inspect the classifications and return to the sources.

679 collected articles can be filtered by brand and topic. 253 have no primary-topic match and need further review; that does not make them irrelevant.

Results

700 records became 679

The workflow removed 19 duplicate URLs and two records that failed the relevance filter.

The remaining 679 records can be searched and reviewed by brand and topic. Of these, 54 have no specific primary-brand match and 253 have no specific primary-topic match, so those records remain available for closer review.

What I learned

Keeping unmatched records visible showed me where the keyword rules needed review. I learned to keep the source links alongside each classification so I could check the original article.

Scope and limits

This is a stored snapshot with keyword classification and lexicon-based polarity. It does not provide live monitoring or validated market sentiment.

Data and method

The data

700 raw RSS records become 679 after removing 19 duplicate URLs and two records failing a keyword filter; classification uses explicit rules.

Method

  1. Collect and retain the raw RSS snapshot; remove 19 duplicate URLs and two records that fail the broad automotive-keyword filter.
  2. Assign a primary brand and topic using transparent keyword dictionaries, while keeping unassigned records visible.
  3. Apply a lexicon-based text-polarity method and build a deterministic overview; retain the underlying records for inspection.
  4. Use the dashboard to filter classifications and return to sources before interpreting the story.

Seven feeds contributed 100 records each to a snapshot collected on 29 June 2026; publication dates span 29 July 2025 to 29 June 2026. Primary assignments and mentions differ: Tesla has seven primary-brand assignments but appears in 40 records.

Project background

The stored raw and processed records retain source context. The feed sample is uneven, and its counts do not measure market or media share.

AI helped refine the code and documentation. The running workflow uses explicit rules, not runtime LLM classification.

Potential use

An automotive researcher following developments across brands and topics.

The decision

Which collected articles are relevant to the question being researched, and which classifications need checking?

Original project outputs
Transparent keyword assignments organise the snapshot. These counts measure collected records, not market or media share.
The original overview dashboard provides entry points into the stored snapshot and its underlying records.
Detailed limitations

Keyword ties follow dictionary order; snippets and multilingual wording can misclassify a record. Lexicon polarity is not consumer or investor sentiment. This is a static, uneven feed sample; no live monitoring, RAG, embeddings or LLM-based analysis is implemented.

Labelled classification checks, broader source coverage, durable publisher links and an assessment of usefulness to analysts.

What I would do next

Hand-label a sample to check classification quality, then improve the topic categories, source URLs and feed coverage.

Explore the repository