vinta/albedo
47.3
Weak · 20 September 2026
3.8k
lines of production code
Scala
with Python
1
measurement over time
What this system is
Albedo is a GitHub repository recommender system that combines a Django web application with Apache Spark for large-scale data processing. It ingests and indexes repository metadata and user interaction history into MySQL and Elasticsearch, then trains multiple recommendation models—including collaborative filtering, content-based, and popularity-based strategies—using Spark MLlib. The system exposes these recommendations through a web interface while providing offline training and evaluation capabilities for various ranking algorithms.
Features
Added development environment configuration and helper scripts
New files have been added to the .docker-assets directory to support local development. This includes environment configuration files for Django (django.env) and MySQL (mysql.env) with default credentials, an Elasticsearch configuration (elasticsearch.yml) for single-node search, and helper scripts for managing the development lifecycle: django\_start.sh for initializing the Python environment, wait\_container.sh for ensuring service availability, and django\_bash\_completion.sh for enabling command-line tab completion for Django management commands.
.docker-assets · high confidence
Albedo project initialized with Django and MySQL
The albedo application has been set up as a new Django 1.11 project. This change introduces the core project structure, including settings, URL routing, and WSGI configuration. A key behavioral aspect is the database configuration, which is now set to use MySQL (host 'mysql', port 3306) instead of the previous SQLite backend, requiring a MySQL instance to be available for the application to function.
albedo · high confidence
Initial project scaffolding for Albedo recommender system
This change introduces the foundational build and deployment infrastructure for the Albedo project, a GitHub repository recommender system built with Apache Spark and Django. It adds a Dockerfile for the Python/Django application, a docker-compose.yml to orchestrate Django, MySQL, and Elasticsearch services, and a comprehensive Makefile with targets for managing local Spark clusters, running Spark jobs (such as ALS and content-based recommenders) locally or on Google Cloud Dataproc, and handling data collection. The commit also includes an IntelliJ IDEA module file defining Scala and Spark dependencies, a .gitignore for the mixed Python/Scala/Django environment, and a detailed README documenting the setup and usage of the various recommendation models.
(repo-wide) · high confidence
Initial release of the GitHub data tracking application
This change introduces the core application structure for tracking GitHub repositories and users. It establishes Django models for repository information, user profiles, user relationships, and star history, along with the corresponding database migrations. The application includes an Elasticsearch mapping for full-text search of repository data, Django admin configurations for managing these entities, and utility scripts for building user-item interaction matrices and timing operations. A basic web view is also provided to serve the initial index page.
app · high confidence
Introduce ALS-based recommendation pipeline with disabled data cleaning
Added Python scripts (\cross\_validate\_als.py\, \train\_als.py\) and supporting toolkit modules (\albedo\_toolkit\) to perform collaborative filtering using Spark's ALS algorithm. The pipeline loads raw data from a local MySQL database, formats it via \RatingBuilder\, and processes predictions with \PredictionProcessor\. Notably, the \DataCleaner\ transformer is defined in the toolkit but explicitly commented out in both the training and cross-validation scripts, meaning the recommendation models are trained and evaluated on raw, unfiltered data without the item/user star-count filtering logic.
src/main/python · high confidence
Introduction of application settings and utility helpers
A new settings package has been added to define core application configuration and utility functions. It exposes configurable paths for data storage and Spark checkpoints (defaulting to ./spark-data and ./spark-data/checkpoint respectively) and provides helper methods to generate date-stamped strings and MD5 hashes, supporting internal data management and processing workflows.
src/main/scala/ws/vinta/albedo/settings · high confidence
New RankingEvaluator for top-k recommendation metrics
A new \RankingEvaluator\ class has been added to the \ws.vinta.albedo.evaluators\ package to evaluate recommendation models using top-k metrics. Users can now assess ranking quality by configuring parameters for the evaluation depth (\k\, default 15), user and item column names, and the specific metric to calculate: NDCG@k, Precision@k, or MAP. The evaluator handles schema validation and computes the selected metric by joining predicted and actual item lists, slicing them to the top-k results, and leveraging Spark's \RankingMetrics\. A companion object provides utility methods to load actual user-item data from Parquet files and transform datasets into the required ranked list format.
src/main/scala/ws/vinta/albedo/evaluators · high confidence
New Spark-based recommendation system components and builders
This change introduces a suite of new Scala entry points and builders in the \ws.vinta.albedo\ package to support a multi-strategy recommendation engine. It adds builders for collaborative filtering (ALS), content-based, curation-based, and popularity-based recommenders, each including logic to load data, generate recommendations, and evaluate results using NDCG. Additionally, it introduces a Logistic Regression ranker that combines user and repository profiles with ALS scores, featuring a cross-validation pipeline for hyperparameter tuning. Supporting infrastructure includes profile builders for users and repositories, a Word2Vec corpus builder for text embeddings, and a playground for model experimentation.
src/main/scala/ws/vinta/albedo · high confidence
New data schemas for user, repository, and recommendation models
The \ws.vinta.albedo.schemas\ package now defines a set of case classes that structure the application's core data models. This includes \UserInfo\ and \RepoInfo\ for detailed user and repository metadata, \Starring\ and \Relation\ for tracking user interactions and social connections, and \PopularRepo\/\UserPopularRepo\ for aggregating repository popularity. Additionally, new classes \Recommendation\, \UserRecommendations\, and \UserItems\ have been introduced to support the recommendation engine, specifically enabling the storage and retrieval of item sequences and predicted ratings for users.
src/main/scala/ws/vinta/albedo/schemas · high confidence
New feature transformers for custom functions and vector assembly
Added two new ML transformers to the feature package: FuncTransformer allows users to apply a custom UserDefinedFunction to map an input column to an output column, supporting serialization for save/load operations; SimpleVectorAssembler merges multiple input columns (numeric, boolean, or vector types) into a single sparse vector column, automatically casting numeric and boolean values to doubles.
src/main/scala/org · high confidence
New management commands for data collection, Elasticsearch sync, and recommendation training
This change introduces a suite of new Django management commands in the \app/management\ module to support data ingestion and recommendation workflows. The \collect\_data\ command allows users to fetch GitHub user and repository information (including followers, following, and starred repos) using provided API tokens, while \drop\_data\ clears this collected relationship and starring data. A new \sync\_data\_to\_es\ command batches and indexes repository metadata into Elasticsearch. Additionally, four new commands (\train\_content\_based\, \train\_graphlab\, \train\_item\_cf\, \train\_user\_cf\) provide offline training capabilities for various recommendation algorithms (content-based, GraphLab, item-based collaborative filtering, and user-based collaborative filtering) to generate repository suggestions for specific users.
app/management · high confidence
New recommender implementations for ALS, content-based, curation, and popularity strategies
The system now includes four new recommendation strategies implemented in the \recommenders\ package: \ALSRecommender\ (collaborative filtering using Spark MLlib), \ContentRecommender\ (content-based filtering via Elasticsearch), \CurationRecommender\ (recommendations based on curated user stars), and \PopularityRecommender\ (recommendations based on repository popularity and recency). These new classes extend the \Recommender\ abstract base class, which defines standard parameters for user/item columns, scoring, and top-K limits, enabling users to select different recommendation sources for their experience.
src/main/scala/ws/vinta/albedo/recommenders · high confidence
New text processing and ranking transformers added
Added five new Spark ML transformers to the library: HanLPTokenizer for Chinese text segmentation with stop-word removal, SnowballStemmer for English word stemming, IntermediateCacher to cache specific DataFrame columns, NegativeBalancer to generate negative samples for recommendation models, and RankingMetricFormatter to prepare predictions for ranking metrics evaluation.
src/main/scala/ws/vinta/albedo/transformers · high confidence
New utility functions for data cleaning and database access
Added a new \closures\ package containing utility objects for data processing. \DBFunctions\ provides a method to query user-starred repositories from a MySQL database. \StringFunctions\ offers regex-based utilities for extracting words (including CJK characters) and email domains. \UDFs\ introduces several Spark User Defined Functions for cleaning company names, email addresses, and locations, as well as for converting Spark vectors to arrays and counting non-zero elements.
src/main/scala/ws/vinta/albedo/closures · high confidence
New utility libraries for data loading, model persistence, and schema validation
The Albedo platform introduces three new utility objects in the utils package to standardize data and model handling. DatasetUtils provides helper methods to load raw data (user info, repo info, starring, and relations) from a MySQL database into Parquet files, automatically creating them if they do not exist, and includes a new randomSplitByUser function for stratified train/test splitting. ModelUtils offers a generic loadOrCreateModel method that loads Spark ML models from disk or creates and saves them if the path is missing. SchemaUtils adds utilities to compare data types ignoring nullability and to validate specific column types against a schema.
src/main/scala/ws/vinta/albedo/utils · high confidence
Test coverage
Added empty test class for Albedo
A new empty test class \AlbedoTest\ has been added in the \src/test/scala/ws/vinta/albedo\ package. This serves as a placeholder for future test coverage of the Albedo component.
src/test · high confidence
Dependencies
Initial dependency manifests for Java/Scala and Python services
Added pom.xml and requirements.txt to define the project's build and runtime dependencies. The Java/Scala side (pom.xml) establishes a Maven build using Java 8 and Scala 2.11, incorporating Apache Spark 2.2.0 (core, SQL, MLlib), Azure's MMLSpark 0.9, Elasticsearch 5.6.2 client, MySQL connector, and NLP libraries (HanLP, Snowball), with the maven-shade-plugin configured to produce an Uber JAR. The Python side (requirements.txt) pins Django 1.11.29, NLTK 3.4.5, requests 2.20.0, and data science libraries (scikit-learn, scipy, pandas, numpy), along with Elasticsearch DSL and Jupyter.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 47.
Lenses
- Code Health 97
- Architecture 97
- Maturity 49
- Readiness 21
- Security 88
Changes since last survey
- 300 commits — 271 feature/other, 29 fixes
By area
- src/main — 211 commits
- (root) — 75 commits
- .idea/runConfigurations — 9 commits
- (repo) — 4 commits
- .idea/artifacts — 1 commit
Notable commits
- fix: fix
- fix: fix
- fix: fix
- fix: fix
- fix: fix
- fix: fix
- fix: fix
- fix: fix
- fix: fix
- fix: fix Makefile
- fix: fix again
- fix: fix als cv
- fix: fix copy
- fix: fix copy() return type
- fix: fix descriptions
- fix: fix encoder errors
- fix: fix java.lang.SecurityException: Invalid signature file digest for Manifest
- fix: fix mvn compile
- fix: fix path
- fix: fix repo profile builder
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
vinta/albedo was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit be94cad3e806616850af985e5befffa2f0898a21 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.