sarrabenyahia/webscrap_health_monitoring
39.3
Weak · 20 September 2026
453
lines of production code
Python
primary language
4
measurements over time
What this system is
This system is a Python-based web scraping tool designed to extract and process the WHO ATC/DDD Index data from the FHI website. It utilizes Scrapy and BeautifulSoup to navigate the hierarchical classification structure, fetching drug codes and definitions across multiple levels. The pipeline then cleans, merges, and formats this data into a unified Excel file for downstream use.
Features
Initial release of the WHO ATC/DDD index scraper
Added a new Scrapy-based web scraper for the WHO Anatomical Therapeutic Chemical (ATC) classification system. The project includes configuration files, middleware, and pipelines, along with two spider implementations: a primary spider (\whocc\_spider.py\) and a multi-level spider (\fourth.py\) that navigates through five levels of the ATC index to extract drug codes and names from the www.whocc.no website.
_whocc\scraper · high confidence
New ATC DDD Index data processing pipeline
Added a new script and supporting module to fetch, parse, and concatenate the WHO ATC DDD Index data from the FHI website into a single Excel file. The pipeline retrieves hierarchical ATC codes (levels 1-5) and associated drug definitions, merges them into a unified dataset, fills missing values using forward/backward fill, sorts the results, and flags entries with valid DDD values.
bs4 · high confidence
Behavioural changes
Project documentation and repository configuration updated
The repository has been rebranded from 'webscrap\_health\monitoring' to 'WHOCC ATC-DDD Index WebScraping' with a comprehensive README detailing the web scraping functionality for the WHOCC ATC DDD Index, including prerequisites (Python 3.11, Pandas, BeautifulSoup, httpx), usage instructions, and file descriptions. Additionally, the .gitignore file was updated to exclude CSV files (\.csv) from version control.
(repo-wide) · high confidence
Dependencies
Initial dependency manifest for web scraping and data processing
Added a new requirements.txt file that pins the project's Python dependencies, including Scrapy and Selenium for web scraping, pandas and openpyxl for data processing, and beautifulsoup4 (bs4) for HTML parsing.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 38 → 39 (+1.4)
- Rubric changed (rubric-2026.08.17 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 98 → 99 (+1.4)
- Architecture 69 → 69 (+0.0)
- Maturity 33 → 33 (+0.0)
- Readiness 15 → 21 (+6.0)
- Security 100 → 80 (-20.0)
Resolved (9)
- Dependency hygiene not measured — no supported dependency manifest was read
- Duplicated block (9–12 lines × 2) (whocc_scraper/whocc_scraper/spiders/fourth.py)
- Duplicated block (9–12 lines × 3) (whocc_scraper/whocc_scraper/spiders/fourth.py)
- No exposed public API
- The README does not mention license or acknowledgements, which are standard metadata for a public project. (README.md)
- early-stage repository — too few commits for a meaningful bus factor
- early-stage repository — too little history to judge knowledge freshness
- git history depth insufficient
- git history depth insufficient
New (104)
- Critical CVE: [GHSA redacted] (requirements.txt)
- Critical CVE: [GHSA redacted] (requirements.txt)
- Documentation: no architecture or design documentation (README.md)
- Documentation: no licence statement (README.md)
- Duplicated block (8–11 lines × 5) (whocc_scraper/whocc_scraper/spiders/fourth.py)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- High CVE: [GHSA redacted] (requirements.txt)
- Medium CVE: [GHSA redacted] (requirements.txt)
- Medium CVE: [GHSA redacted] (requirements.txt)
- Medium CVE: [GHSA redacted] (requirements.txt)
- …and 84 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
sarrabenyahia/webscrap_health_monitoring was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 2e0f54630c80ed016359e48956337789c1b226ca — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-28e75b8e3254.