03 / Research Automation & Information Analysis
Automotive News Intelligence Workflow
Rule-based news research and source review
The problem
I wanted an easier way to follow relevant automotive and EV developments without getting lost in the volume of news.
Personal portfolio case · Stored RSS snapshot
Source & implementationWhy I built it
Following automotive news made it difficult to separate useful developments from noise. I explored whether collecting articles, removing duplicates and organising them by brand and topic would make the material easier to review.
I wanted to search by brand or topic and open the original articles when a result needed checking.
My contribution
I built a workflow around a stored RSS snapshot: collect records, remove duplicate URLs, apply a relevance filter and group the remaining records with keyword rules. The dashboard lets me inspect the classifications and return to the sources.
Results
700 records became 679
The workflow removed 19 duplicate URLs and two records that failed the relevance filter.
The remaining 679 records can be searched and reviewed by brand and topic. Of these, 54 have no specific primary-brand match and 253 have no specific primary-topic match, so those records remain available for closer review.
What I learned
Keeping unmatched records visible showed me where the keyword rules needed review. I learned to keep the source links alongside each classification so I could check the original article.
Scope and limits
This is a stored snapshot with keyword classification and lexicon-based polarity. It does not provide live monitoring or validated market sentiment.
Data and method
The data
700 raw RSS records become 679 after removing 19 duplicate URLs and two records failing a keyword filter; classification uses explicit rules.
Method
- Collect and retain the raw RSS snapshot; remove 19 duplicate URLs and two records that fail the broad automotive-keyword filter.
- Assign a primary brand and topic using transparent keyword dictionaries, while keeping unassigned records visible.
- Apply a lexicon-based text-polarity method and build a deterministic overview; retain the underlying records for inspection.
- Use the dashboard to filter classifications and return to sources before interpreting the story.
Seven feeds contributed 100 records each to a snapshot collected on 29 June 2026; publication dates span 29 July 2025 to 29 June 2026. Primary assignments and mentions differ: Tesla has seven primary-brand assignments but appears in 40 records.
Project background
The stored raw and processed records retain source context. The feed sample is uneven, and its counts do not measure market or media share.
AI helped refine the code and documentation. The running workflow uses explicit rules, not runtime LLM classification.
Potential use
An automotive researcher following developments across brands and topics.
The decision
Which collected articles are relevant to the question being researched, and which classifications need checking?
Original project outputs
Detailed limitations
Keyword ties follow dictionary order; snippets and multilingual wording can misclassify a record. Lexicon polarity is not consumer or investor sentiment. This is a static, uneven feed sample; no live monitoring, RAG, embeddings or LLM-based analysis is implemented.
Labelled classification checks, broader source coverage, durable publisher links and an assessment of usefulness to analysts.