Michael-A-Kuykendall/shimmy
63.9
Adequate · 29 September 2026
13.2k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is a Rust-based inference server that serves GGUF models via an OpenAI-compatible API, utilizing the Airframe WebGPU engine as its default backend. It features automatic model discovery from local directories and Ollama, unified Jinja-based prompt rendering, and strict input validation. The codebase has been refactored to consolidate testing and remove legacy backends, focusing on efficient, template-aware generation for chat and completion endpoints.
Features
Introduce automatic GGUF model discovery across common directories and Ollama
The application now automatically scans for GGUF model files in standard locations (such as Hugging Face cache, LM Studio, and local model folders) and integrates with Ollama by parsing its manifest and blob structure. This feature populates the model list without manual configuration, supporting both standard GGUF files and Ollama-specific model formats, while handling sharded models and respecting environment variables for custom search paths.
_src/auto\discovery · high confidence
Unified Jinja-first prompt rendering and OpenAPI documentation
Prompt construction now uses a single renderer that prioritizes the model's native GGUF Jinja chat template, falling back to hardcoded family templates or raw text, which improves compatibility with diverse model architectures. The API layer has been updated to use this unified rendering for both chat and completion endpoints, and a new \raw\_prompt\ option allows bypassing templates for base models. Additionally, an embedded OpenAPI specification and Swagger UI documentation page are now available to help users explore the API endpoints.
src · high confidence
Removals
Removed Vision testing Dockerfiles for cross-platform builds
The Dockerfiles used for cross-compilation and build verification of the Vision feature on Linux ARM64, Linux CUDA, macOS, and Windows have been removed from the packaging directory. This change excises the Vision product from the public release branch, meaning these specific CI/CD build targets are no longer available for testing or distribution.
packaging · high confidence
Behavioural changes
Airframe GPU engine replaces legacy backends as the default inference path
The engine module has been refactored to use the new Airframe GPU engine as the primary inference backend, replacing the previous HuggingFace, Llama, and MLX engines which have been removed. The \InferenceEngineAdapter\ now instantiates \AirframeEngine\ (when the \airframe\ feature is enabled) and \SafeTensorsEngine\, routing GGUF files to Airframe by default. The \GenOptions\ struct has been extended with new fields (\grammar\_mode\, \fse\_reject\_patterns\, \math\_bypass\, \trace\_path\, \session\_id\, \raw\_prompt\) to support advanced generation controls, and the \ModelSpec\ now includes a \chat\_template\ field to enable Jinja-first prompt rendering, ensuring instruct models are correctly templated before generation.
src/engine · high confidence
OpenAI-compatible API input validation and Ollama model discovery
The OpenAI-compatible API layer now enforces stricter input validation by rejecting requests with empty message arrays and explicitly blocking unsupported multimodal content (returning a 400 Bad Request with an 'unsupported\_media' error code). Additionally, the module introduces an Ollama-compatible GET /api/tags endpoint, allowing clients like AnythingLLM, SillyTavern, Zed, and Open WebUI to discover available models via this route. The codebase has also been refactored to split type definitions into a dedicated \types.rs\ module.
_src/openai\compat · high confidence
Fixes
Consolidated regression test suite and removed obsolete benchmark scripts
The regression testing infrastructure has been consolidated from a large, auto-discovered set of individual test files into three main test targets: core, handlers, and compile\_checks. The automated runner scripts (run-regression-tests.sh and run-regression-tests-auto.sh) have been updated to execute these consolidated suites, reflecting changes such as the addition of SSE regression tests, fixes for template compilation, and updates to error handling and API compatibility checks. Additionally, the standalone Python benchmark and MoE stress test scripts (benchmark.py and moe\_stress\_test.py) have been removed from the repository.
scripts · high confidence
Test coverage
Consolidated test suite with new compile-time, core, routing, and handler coverage; Regression test suite consolidated and modernized for Airframe engine; Updated PPT contract tests to target Airframe backend.
Dependencies
Shimmy v2.6.4: Airframe GPU engine default and dependency refresh
Shimmy has been updated to version 2.6.4, making the Airframe native WebGPU engine the default inference backend (previously, the default included llama.cpp and HuggingFace features). The \airframe\ dependency is now pinned to version 0.4.3, and the \shimmyjinja\ dependency is added at version 0.5.0. Several underlying dependencies have been updated, including \reqwest\ (0.11 to 0.12), \anyhow\ (1.0.100 to 1.0.104), \async-trait\ (0.1.89 to 0.1.92), and \assert\_cmd\ (2.1.1 to 2.2.2). Legacy llama.cpp feature flags remain as deprecated stubs, and the package description now highlights cross-platform WebGPU acceleration via Airframe.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 67 → 64 (-2.9)
- Rubric changed (rubric-2026.09.10 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 92 → 92 (+0.0)
- Architecture 100 → 97 (-3.1)
- Maturity 67 → 67 (+0.2)
- Readiness 56 → 54 (-2.1)
- Security 73 → 76 (+2.9)
- Performance 71 (new)
Resolved (6)
- Dependency hygiene PARTLY measured — npm pinning read, dependency currency not (no pnpm-resolved versions to grade)
- Documentation: no installation or build instructions (README.md)
- Hotspot: src/api.rs (src/api.rs)
- Hotspot: src/openai_compat/mod.rs (src/openai_compat/mod.rs)
- Hotspot: src/prompt_render.rs (src/prompt_render.rs)
- Hotspot: src/templates.rs (src/templates.rs)
New (12)
- Duplicate message types for different protocols. While distinct protocols require distinct types, the internal structure varies significantly (e.g., ChatMessage.content is String, OAIMessage.content is MessageContent, AnthropicMessage.content is AnthropicContent). This makes it difficult to write protocol-agnostic code or convert between them.
- Duplicate types with inconsistent field types. auto_discovery.DiscoveredModel and discovery.DiscoveredModel appear to be duplicate discovery result types. Furthermore, the size_bytes field is u64 in one and String in the others, creating serialization and usage friction.
- Duplicated block (7 lines × 2) (templates/frameworks/fastapi/main.py)
- High: security finding (details withheld)
- Inconsistent context length representation. Cli and ModelEntry use String for ctx/ctx_len, while ModelSpec uses usize. This inconsistency propagates through the configuration and model loading pipeline.
- Inconsistent numeric representation. parameter_count is represented as String in all discovery/API types, which is unusual for numeric metadata and suggests a lack of standardization or potential precision loss issues compared to using u64 or f64.
- Inconsistent path types. lora_path is String in multiple places, whereas path fields in discovery types use PathBuf. This forces users to convert between String and PathBuf when moving between discovery, registry, and engine specs.
- Massive inconsistency in parameter types. API request types (GenerateRequest, ChatCompletionRequest, CompletionRequest) and Cache keys use String for numeric/boolean parameters (temperature, max_tokens, stream, etc.), while the internal engine options (GenOptions) use native types (f32, i32, bool). This forces unnecessary parsing at the API boundary and serialization at the engine boundary.
- Medium vulnerability: RUSTSEC-2026-0285 (Cargo.lock)
- Outdated: clap
- Outdated: uuid
- Test suite declares a test script but contains no test files: templates/frameworks/express/
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Michael-A-Kuykendall/shimmy was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 4895730a700cb164ac2855d2ad35cb3ce59e557e — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5ff527f25b99.