Skip to content
CAI
Software that uses CAICheck a score

lucidrains/vit-pytorch

51.8

Adequate · 18 September 2026

29.8k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Python library for implementing and training Vision Transformer (ViT) models and their variants. It provides a collection of architectural modules, including CaiT, CCT, and CrossViT, along with utilities for self-supervised learning and video processing. The library supports distributed training via Accelerate and includes example notebooks and scripts for common datasets like CIFAR-100 and Cats & Dogs.

Features

Added training script for ViT with decorrelation auxiliary loss

A new training script (train\_vit\_decorr.py) is provided to demonstrate training a Vision Transformer with decorrelation auxiliary losses. This script uses the accelerate library for mixed-precision and distributed training, wandb for experiment tracking, and CIFAR-100 as the dataset. It also includes a MANIFEST.in file to ensure tests are included in the package distribution.

(repo-wide) · high confidence

New Vision Transformer Architectures and Video Processing Capabilities

This release introduces a wide range of new Vision Transformer (ViT) variants and specialized modules, including CaiT, CCT, CrossViT, CrossFormer, CvT, DeepViT, and Adaptive Token Sampling (ATS) ViT. It also adds self-supervised learning implementations for DINO and MAE, a video processing wrapper (AcceptVideoWrapper) for handling temporal sequences, and a 3D CCT variant. These additions expand the library's support for diverse transformer architectures and multi-modal inputs.

_vit\pytorch · high confidence

New example notebook for Visual Transformer on Cats & Dogs dataset

Added a new Jupyter notebook example (\examples/cats\_and\_dogs.ipynb\) demonstrating how to train a Visual Transformer using Linformer on the Dogs vs. Cats dataset. The notebook includes setup for dependencies (\vit\_pytorch\, \linformer\), data loading, and model training configuration.

examples · high confidence

Test coverage

Added tests for DINO loss stability and ViT model functionality

Added new test coverage for the DINO distillation loss function, specifically verifying that loss values and gradients remain finite and numerically stable when using low-precision data types (float16 and bfloat16), and confirming gradient correctness against a reference implementation. Additionally, added a basic functional test for the ViT model to ensure it produces correctly shaped output logits.

tests · high confidence

Dependencies

vit-pytorch version 1.26.6 release

The package has been updated to version 1.26.6. The build system now requires setuptools\>=61 and wheel, and the project specifies a minimum Python version of 3.8. Core dependencies include einops\>=0.8.2 and torch\>=2.5, with torchvision as a dependency. Test dependencies are pinned to pytest, torch==2.5.0, and torchvision==0.20.0.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 52.

Lenses

  • Code Health 100
  • Architecture 100
  • Maturity 36
  • Readiness 33
  • Security 100

Changes since last survey

  • 300 commits — 267 feature/other, 33 fixes

By area

  • (root) — 248 commits
  • (repo) — 13 commits
  • vit_pytorch/mae.py — 5 commits
  • .github/workflows — 4 commits
  • vit_pytorch/vit.py — 3 commits
  • vit_pytorch/crossformer.py — 2 commits
  • vit_pytorch/mobile_vit.py — 2 commits
  • vit_pytorch/na_vit.py — 2 commits
  • vit_pytorch/vivit.py — 2 commits
  • .github/FUNDING.yml — 1 commit
  • examples/cats_and_dogs.ipynb — 1 commit
  • images/mbvit.png — 1 commit
  • images/xcit.png — 1 commit
  • tests/test_dino_loss.py — 1 commit
  • tests/test_vit_with_decorr.py — 1 commit
  • vit_pytorch/cct.py — 1 commit
  • vit_pytorch/cct_3d.py — 1 commit
  • vit_pytorch/mpp.py — 1 commit
  • vit_pytorch/na_vit_nested_tensor.py — 1 commit
  • vit_pytorch/na_vit_nested_tensor_3d.py — 1 commit

Notable commits

  • fix: Fix ViViT Transformer not passing use_flash_attn to Attention and duplicate mask reshape (#360)
  • fix: Merge pull request #101 from zankner/mpp-fix
  • fix: Merge pull request #153 from developer0hye/fix-example
  • fix: double down on dual patch norm, fix MAE and Simmim to be compatible with dual patchnorm
  • fix: fix block repeats in readme example for Nest
  • fix: fix distill
  • fix: fix feature maps in Nest, thanks to @MarkYangjiayi
  • fix: fix hard distillation, thanks to @CiaoHe
  • fix: fix hidden dimension in MaxViT thanks to @arquolo
  • fix: fix linear head in simple vit, thanks to @atkos
  • fix: fix max pool in nest
  • fix: fix maxvit - need feedforwards after attention
  • fix: fix mbconv residual block
  • fix: fix mpp
  • fix: fix mpp
  • fix: fix multiheaded qk rmsnorm in nViT
  • fix: fix positional embed for mean pool case and cleanup
  • fix: fix pypi
  • fix: fix recorder in data parallel situation
  • fix: fix small bug
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

lucidrains/vit-pytorch was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit d01bbb31ec889a338f6048c273dcf2af1a097d30 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.