Skip to content
CAI
Software that uses CAICheck a score

sarrabenyahia/webscrap_health_monitoring

39.3

Weak · 20 September 2026

453

lines of production code

Python

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a Python-based web scraping tool designed to extract and process the WHO ATC/DDD Index data from the FHI website. It utilizes Scrapy and BeautifulSoup to navigate the hierarchical classification structure, fetching drug codes and definitions across multiple levels. The pipeline then cleans, merges, and formats this data into a unified Excel file for downstream use.

Features

Initial release of the WHO ATC/DDD index scraper

Added a new Scrapy-based web scraper for the WHO Anatomical Therapeutic Chemical (ATC) classification system. The project includes configuration files, middleware, and pipelines, along with two spider implementations: a primary spider (\whocc\_spider.py\) and a multi-level spider (\fourth.py\) that navigates through five levels of the ATC index to extract drug codes and names from the www.whocc.no website.

_whocc\scraper · high confidence

New ATC DDD Index data processing pipeline

Added a new script and supporting module to fetch, parse, and concatenate the WHO ATC DDD Index data from the FHI website into a single Excel file. The pipeline retrieves hierarchical ATC codes (levels 1-5) and associated drug definitions, merges them into a unified dataset, fills missing values using forward/backward fill, sorts the results, and flags entries with valid DDD values.

bs4 · high confidence

Behavioural changes

Project documentation and repository configuration updated

The repository has been rebranded from 'webscrap\_health\monitoring' to 'WHOCC ATC-DDD Index WebScraping' with a comprehensive README detailing the web scraping functionality for the WHOCC ATC DDD Index, including prerequisites (Python 3.11, Pandas, BeautifulSoup, httpx), usage instructions, and file descriptions. Additionally, the .gitignore file was updated to exclude CSV files (\.csv) from version control.

(repo-wide) · high confidence

Dependencies

Initial dependency manifest for web scraping and data processing

Added a new requirements.txt file that pins the project's Python dependencies, including Scrapy and Selenium for web scraping, pandas and openpyxl for data processing, and beautifulsoup4 (bs4) for HTML parsing.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 38 → 39 (+1.4)
  • Rubric changed (rubric-2026.08.17 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 98 → 99 (+1.4)
  • Architecture 69 → 69 (+0.0)
  • Maturity 33 → 33 (+0.0)
  • Readiness 15 → 21 (+6.0)
  • Security 100 → 80 (-20.0)

Resolved (9)

  • Dependency hygiene not measured — no supported dependency manifest was read
  • Duplicated block (9–12 lines × 2) (whocc_scraper/whocc_scraper/spiders/fourth.py)
  • Duplicated block (9–12 lines × 3) (whocc_scraper/whocc_scraper/spiders/fourth.py)
  • No exposed public API
  • The README does not mention license or acknowledgements, which are standard metadata for a public project. (README.md)
  • early-stage repository — too few commits for a meaningful bus factor
  • early-stage repository — too little history to judge knowledge freshness
  • git history depth insufficient
  • git history depth insufficient

New (104)

  • Critical CVE: [GHSA redacted] (requirements.txt)
  • Critical CVE: [GHSA redacted] (requirements.txt)
  • Documentation: no architecture or design documentation (README.md)
  • Documentation: no licence statement (README.md)
  • Duplicated block (8–11 lines × 5) (whocc_scraper/whocc_scraper/spiders/fourth.py)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • High CVE: [GHSA redacted] (requirements.txt)
  • Medium CVE: [GHSA redacted] (requirements.txt)
  • Medium CVE: [GHSA redacted] (requirements.txt)
  • Medium CVE: [GHSA redacted] (requirements.txt)
  • …and 84 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

sarrabenyahia/webscrap_health_monitoring was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 2e0f54630c80ed016359e48956337789c1b226ca — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-28e75b8e3254.