Skip to content
CAI
Software that uses CAICheck a score

sculab/GeneMiner2

42.7

Weak · 1 October 2026

13.5k

lines of production code

VB.NET

with Python, C, C++, Haxe

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a bioinformatics pipeline for genome assembly and sequence analysis, featuring tools for filtering, alignment, and consensus generation. It processes genomic data with configurable depth controls and supports compressed binary formats to optimize performance on slower storage. The software includes a graphical interface for configuration and provides documentation and scripts for Linux and Windows environments.

Behavioural changes

Filter performance improvements and binary format support

The filter tool now handles larger data chunks (increased from 64KB to 192KB) and supports a binary output format for faster processing on slower disks, alongside internal optimizations to the hash table and memory management to reduce issues on slower storage. The filter also now correctly sets an exit code when it fails, and includes various bug fixes and reorganized query loops for better stability.

scripts/filter · high confidence

Refined depth controls, new .gm2 support, and alignment fixes

The pipeline now distinguishes between k-mer frequency and read depth, introducing new command-line options: \--depth-low-water-mark\ and \--depth-limit\ for re-filtering, and \--min-coverage\ for assembly to ensure contigs have sufficient read support. The refiltering step now supports the compressed \.gm2\ format, and the assembler includes a new Cython extension (\main\_refilter\_ext.pyx\) to improve performance on slower disks. Additionally, the alignment cleanup in \fix\_alignment.py\ has been corrected to properly handle gaps, and the consensus script now reliably writes error logs to the correct path on Linux.

scripts · high confidence

Updated demo documentation and removed sample data files

The demo documentation has been revised: DEMO1.md now includes a 'Calculate parameters' section, DEMO2.md clarifies memory requirements for genome assembly (advising 4GB) and removes the Plant Mitochondrial Genome instructions, and DEMO3.md adds a warning to avoid Chinese characters in output paths. Additionally, several large Phytozome CDS FASTA files (Dcarota, Eanacua, Hannuus) have been removed from the DEMO3 data directory.

DEMO · high confidence

Version 2.5 release with UI terminology updates and new filtering workflow

The application has been updated to version 2.5 (build 20260414). The user interface terminology in the Calculate and Combine configuration dialogs has been refined: 'Maximum Difference' and 'Max Diff.' are now labeled 'Cluster difference' (or '聚类差异限值' in Chinese), and 'Number of sequences' is now 'Number of samples' (or '最小样本数量'). A new 'Consensus Sequence' configuration form has been added. The filtering process on Windows has been redesigned to use a dedicated 'main\_refilter\_new.exe' tool with a '--use-gm2-format' flag instead of manual file copying, and the final analysis completion message now directs users to check the 'best\_refs' folder.

CODE · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 44 → 43 (-1.6)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 64 → 64 (-0.5)
  • Architecture 100 → 100 (+0.0)
  • Maturity 60 → 60 (+0.0)
  • Readiness 15 → 15 (+0.0)
  • Security 100 → 80 (-20.0)

Resolved (9)

  • Documentation: no architecture or design documentation (README.md)
  • Documentation: no installation or build instructions (README.md)
  • Edited copy of a member (7 corresponding lines) (CODE/Config_Combine.vb)
  • Members sharing a duplicated core (13 members, 50+ identical tokens) (CODE/Main_Form.vb)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (CODE/Config_CP.vb)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (CODE/Config_Dated.vb)
  • Members sharing a duplicated core (7 members, 50+ identical tokens) (CODE/Main_Form.vb)
  • Members sharing a duplicated core (8 members, 50+ identical tokens) (CODE/Main_Form.vb)
  • Members sharing a duplicated core (9 members, 50+ identical tokens) (CODE/Main_Form.vb)

New (16)

  • Duplicated block (105–108 lines × 2) (scripts/PPD.py)
  • Duplicated block (11 lines × 2) (scripts/main_assembler.py)
  • Duplicated block (11 lines × 2) (scripts/split_genes.py)
  • Duplicated block (5 lines × 2) (scripts/main_refilter_new.py)
  • Duplicated block (6 lines × 2) (scripts/build_gb.py)
  • Duplicated block (8 lines × 2) (scripts/build_fq.py)
  • Duplicated block (8 lines × 2) (scripts/unix_command.py)
  • Duplicated block (8 lines × 4) (scripts/PPD.py)
  • Edited copy of a member (7 corresponding lines) (CODE/Config_Combine.vb)
  • End-of-life runtime: .NET net6.0
  • Members sharing a duplicated core (13 members, 50+ identical tokens) (CODE/Main_Form.vb)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (CODE/Config_CP.vb)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (CODE/Config_Dated.vb)
  • Members sharing a duplicated core (7 members, 50+ identical tokens) (CODE/Main_Form.vb)
  • Members sharing a duplicated core (8 members, 50+ identical tokens) (CODE/Main_Form.vb)
  • Members sharing a duplicated core (9 members, 50+ identical tokens) (CODE/Main_Form.vb)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

sculab/GeneMiner2 was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 1 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit f02313f20a97e05fc625a41d28f4b18fcfa013a8 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-e569280dd5e2.