Skip to content
CAI
Software that uses CAICheck a score

JohnSnowLabs/spark-nlp

65.7

Adequate · 20 July 2026

222.7k

lines of production code

Scala

with Python

1

measurement over time

CAI band scale
CAI lens gauges

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 66.

Lenses

  • Code Health 99
  • Architecture 96
  • Maturity 72
  • Readiness 46
  • Security 87
  • Domain Modelling 100

Changes since last survey

  • 300 commits — 273 feature/other, 27 fixes

By area

  • src/main — 87 commits
  • (root) — 36 commits
  • examples/python — 29 commits
  • python/sparknlp — 25 commits
  • src/test — 24 commits
  • docs/api — 22 commits
  • python/test — 18 commits
  • docs/en — 15 commits
  • docs/_posts — 14 commits
  • .github/workflows — 10 commits
  • conda/meta.yaml — 5 commits
  • project/Dependencies.scala — 4 commits
  • docs/_data — 2 commits
  • scripts/colab_setup.sh — 2 commits
  • (repo) — 1 commit
  • .github/ISSUE_TEMPLATE — 1 commit
  • docs/Gemfile — 1 commit
  • docs/_config.yml — 1 commit
  • docs/_frontend — 1 commit
  • docs/assets — 1 commit

Notable commits

  • fix: Adding perfered engine to pretrained and minor fixes (#14689)
  • fix: AutoGGUFModel: fix type annotations on python side
  • fix: AutoGGUFVisionModel notebook minor fixes [skip test]
  • fix: Fix Features test
  • fix: Fix ReaderAssemblerTest PDF test failing
  • fix: Fix Spark NLP example notebooks (#14620)
  • fix: Fix default pretrained gemma3 model
  • fix: Fix google colab link on Readers Notebooks
  • fix: Fix out of memory error when copying big models to a cloud storage
  • fix: Fix python tests
  • fix: Fix tests for LightPipeline bug fix
  • fix: LightPipeline with IDs: Fix Scala Map metadata
  • fix: Revert "Updating class of skipPerferredEngine in scala"
  • fix: Revert "setting skipPerferredEngine to True in AutoGGUFModel and AutoGGUFVisionModel"
  • fix: [SPARKNLP-1174] Fix reading as text file content
  • fix: [SPARKNLP-1260] Fix python test files path
  • fix: [SPARKNLP-1280] Python side: fix pretrained model AutoGGUFVision
  • fix: [SPARKNLP-1297] AutoGGUF python tests fix warning
  • fix: [SPARKNLP-1297] NerDLGraphChecker: Fix tests for pyspark 3.3
  • fix: [SPARKNLP-1309] Fix repeating tokens in WordEmbeddings
  • …and 280 more

Architecture

  • 0 containers · 2 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

JohnSnowLabs/spark-nlp was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 July 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit fcede1de1962b215026c1757d7c054e6d99bb9fe — the exact code this score is about.
  • Scored under rubric-2026.08.17 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer latest.