VirtusLab/orca
68.4
Adequate · 20 September 2026
33k
lines of production code
Scala
primary language
1
measurement over time
What this system is
Orca is a CLI-based orchestration system for automating software development workflows using multiple LLM backends. It executes durable, resumable flows that handle autonomous planning, code implementation, and pull request automation. The system supports interactive agent sessions with human-in-the-loop clarification and integrates with various AI providers to manage the full lifecycle of coding tasks.
Features
Add Gemini CLI as an LLM backend
Users can now select the Gemini CLI as an LLM backend for agent runs. This change introduces the \GeminiBackend\ and \GeminiConversation\ components, which drive the \gemini\ CLI via stdio using a \stream-json\ protocol. The implementation supports both autonomous and interactive modes, including session resumption via \--resume\ and an MCP \ask\_user\ bridge for interactive tool approvals. The backend exposes two model variants: \gemini-2.5-pro\ (default) and \gemini-2.5-flash\ (via the \.flash\ accessor), with structured output handled via raw text prompt enforcement.
gemini/src/main · high confidence
Initial project scaffolding and documentation
The repository is bootstrapped with a multi-module sbt build structure (tools, flow, runner, and backend modules for Claude, Codex, Gemini, Opencode, and Pi), along with essential project documentation including AGENTS.md for internal conventions, CONTRIBUTING.md for build and test instructions, and a comprehensive README.md detailing installation, usage, and flow scripting. The project is licensed under Apache 2.0, includes an installation script for the \orca\ CLI, and configures Scala 3.9.0 with scalafmt 3.8.3 (excluding capture-checking syntax files).
(repo-wide) · high confidence
Introduce Codex backend with autonomous and interactive modes
Users can now run agents via the Codex CLI backend, which supports both autonomous execution and interactive sessions. The default configuration uses the GPT-5.6 Sol model, with an option to switch to the cheaper GPT-5.4 Mini model for lightweight tasks. Interactive sessions include an MCP-based 'ask\_user' bridge that allows the agent to surface clarifying questions to the user, while autonomous runs operate without this overhead. Session state is managed server-side, allowing subsequent turns to resume existing threads using the same client-assigned session ID.
codex/src/main · high confidence
Introduce OpenCode as a third LLM tool backend
Added OpenCode as a new LLM tool backend, allowing users to drive agents via a shared local \opencode serve\ process over HTTP and SSE. This includes a JDK-based HTTP client for server communication, a launcher abstraction that supports default execution or Ollama-wrapped local models, and provider-prefixed model accessors (Anthropic and OpenAI) with automatic selection of a cost-effective model for incidental work.
opencode/src/main · high confidence
Introduce Pi agent backend with user-interaction support
Added a new Pi agent backend that enables the system to interact with Pi-compatible AI models. This includes a default agent implementation, wiring for standard configuration, and a temporary extension mechanism that dynamically loads an 'Ask User' tool. This tool allows the agent to pause execution and request concise clarifications from the human user, with automatic cleanup of temporary extension files upon session close.
pi/src/main · high confidence
Introduce durable, resumable flow stages and sessions
The flow module now supports durable, resumable execution through a new stage-based model. Runs are broken into named stages that automatically persist their results to a JSON progress log; on restart, completed stages are skipped and their results replayed, while failed stages re-run from their recorded baseline. LLM sessions are now durable and keyed by name and detail, allowing them to survive crashes and resume seamlessly. The system also introduces role-based agent selection (planning, coding, review) resolved from project settings, and enforces strict thread affinity to prevent race conditions during stage and session management.
flow/src/main/scala/orca · high confidence
Introduce new Claude agent implementation with model-specific variants
This change adds the core Scala classes for the new Claude agent backend (\ClaudeAgents\, \ClaudeConversation\, \DefaultClaudeAgent\). It establishes the default agent configuration using the Opus-1M model for coding tasks, while exposing specific model aliases (Haiku, Sonnet, Fable) for different use cases. The implementation includes a conversation driver that handles stream-json protocol translation, manages tool calls (including \ask\_user\ and structured output), and supports network tool allowlists.
claude/src/main · high confidence
Introduce project-specific code review and build configuration
The .orca tool now supports project-specific reviewers and settings, allowing teams to enforce custom conventions and tooling without relying solely on global defaults. A new \.orca/reviewers/orca.md\ file defines repository-specific review rules covering code style, state management, and modeling, while \.orca/settings.properties\ configures local format, lint, and test commands (e.g., \sbt scalafmt\, \sbt compile\, \sbt test\). This enables granular control over how the tool interacts with the project's build and review processes.
.orca · high confidence
Introduce the Orca interactive shell and CLI interface
This change introduces the \orca\ shell executable, providing both an interactive menu-driven interface and a non-interactive command-line interface (CLI) for managing and running flows. The interactive shell features a main menu with options to run, view, edit, create, and fork flows, as well as reconfigure agents and clear stack settings. It includes a first-run wizard for initial configuration and supports resuming interrupted runs and harness sessions. The CLI exposes commands like \orca run\, \orca create\, \orca fork\, \orca config\, and \orca continue\ for scripted or direct usage. Additionally, a minimal Scala syntax highlighting definition (\scala.nanorc\) is added for use within the shell's terminal interface.
shell/src/main · high confidence
New Claude and Codex LLM backends with MCP support
Added dedicated backends for the Claude Code and Codex CLI tools, enabling both autonomous and interactive agent runs. The Claude backend uses the \stream-json\ protocol to parse structured events and supports host-served MCP servers (including an \ask\_user\ bridge) for interactive sessions. The Codex backend drives \codex exec\ via JSONL, handling multi-turn resumption and sandbox enforcement. Both backends map agent configurations to specific CLI flags, enforce tool permissions (read-only, network-only, full), and integrate with the existing session and event systems.
repository · high confidence
New PR automation helpers and review findings integration
This change introduces a new set of utilities in the \orca.pr\ package to streamline Pull Request creation and handling. It adds \openPrFromBranch\ to automate the sequence of pushing a branch, generating a PR summary via an AI agent, and creating the PR, while \bodyWithOpenFindings\ allows flows to append a section to the PR body listing any review issues that were left open or declined by the fixer. Supporting infrastructure includes \PrPrompts\ for loading prompt templates, \orcaCommentMarker\ for stable GitHub comment identification, and \recordOpenedPr\ to persist the opened PR URL as published work for resume safety.
flow/src/main/scala/orca/pr · high confidence
New and updated flow scripts for autonomous coding and review
The \flows\ directory now includes several new and updated Scala CLI scripts that define how the Orca agent operates. \implement-enhanced.sc\ adds an autonomous planning flow that includes self-review of the plan and a dedicated documentation update stage before opening a PR. \implement-interactive.sc\ introduces an interactive planning mode where the agent can ask clarifying questions before generating a plan. \issue-pr-bugfix.sc\ provides a bug-fix workflow that requires a CI-verified failing test before implementing a fix. \review.sc\ is a new read-only flow for reviewing PRs, branches, or local changes without making fixes. Additionally, \implement.sc\, \issue-pr.sc\, and \simple.sc\ have been updated to use the \orca\ library version 0.1.7 and Scala 3.9.0, and now support GitHub URLs in issue/PR references alongside short pointers.
flows · high confidence
New runnable examples for autonomous and interactive planning flows
Added two new runnable examples in the \examples/runnable\ directory: \01-simple\ demonstrates a one-shot autonomous planning flow where the agent plans and implements tasks in memory without resume capability, while \02-interactive\ extends this shape by allowing the planner to ask clarifying questions via an \ask\_user\ MCP tool before producing a plan. Both examples include seed scripts (\create-test-project.sh\) that set up a small Rust calculator crate, initialize a git repository, and optionally pin the Orca dependency to a local build or seed stack settings for deterministic runs.
examples · high confidence
Structured output prompts and capability-split concurrency primitives
The tools module now provides dedicated prompt templates for autonomous and interactive structured-output flows (including retry handling for raw JSON and tool-based outputs) and introduces a capability-split concurrency model using capture checking. This model enforces that workspace mutations (WorkspaceWrite) remain exclusive to the flow thread while allowing LLM calls (InStage) to cross fork boundaries, ensuring safe parallel fan-outs in flow scripts without compromising write integrity.
tools/src/main · high confidence
Behavioural changes
Centralized cost resolution and hardened event listener dispatch
The event system now ensures consistent cost reporting by introducing a CostResolvingDispatcher that calculates token costs once at the entry point, preventing discrepancies between the terminal summary, cost logs, and user listeners. Additionally, the EventDispatcher has been hardened to improve reliability: listeners that throw exceptions are now automatically quarantined for the remainder of the run, and their failures are logged to both the trace file and stderr to ensure visibility even if the primary logging system is affected.
flow/src/main/scala/orca/events · high confidence
New RunTarget model and consolidated exports for flow scripts
The runner introduces a new \RunTarget\ enum to explicitly model how a run handles uncommitted files (stash vs. keep) and where it commits (new branch, current branch, or worktree), enforcing mutually exclusive flag combinations like \--worktree\ with \--skip-branch\ or \--keep-changes\ at the type level. Alongside this, a new \exports.scala\ file consolidates re-exports of the public API—including agents, events, planning, PR handling, review loops, and tool outcomes—allowing flow scripts to import the entire surface with a single \import orca.{\*, given}\ statement.
runner/src/main/scala/orca · high confidence
New review loop formatting, logging, and prompt infrastructure
The review loop now uses dedicated formatting and logging components to improve output clarity and debuggability. Review outcomes are rendered with consistent wrapping and indentation, explicitly showing issue keys, file locations, and suggestions, while open findings are listed with reasons for being unresolved. A new logging module records detailed debug traces of all review turns (initial, re-review, fix, and picker selections), including the specific change set shape and content sent to agents, without cluttering the console. Prompt templates are now loaded from classpath resources, allowing for structured, reusable instructions for initial reviews, re-reviews, and reviewer selection, with support for task context, user requests, and diff baselines.
flow/src/main/scala/orca/review · high confidence
Redesigned terminal output with stage indentation, status bar, and tool step summaries
The terminal interface now features a structured event log with stage-based indentation, a persistent animated status bar at the bottom, and concise one-line summaries for tool calls. Read-only tool invocations are collapsed into a single repeated line to reduce noise, while parallel agent outputs are distinguished by dark-gray attribution prefixes. The renderer enforces strict line budgets to prevent wrapping, handles multiline input via JLine with explicit newline support, and forces UTF-8 encoding on stderr to ensure non-ASCII glyphs display correctly across all locales.
runner/src/main/scala/orca/runner/terminal · high confidence
Refactor planning prompts and triage outcomes into dedicated types
The planning module now loads prompt instructions (Planning, AssessThenPlan, Triage, Review) from classpath resources via a new PlanPrompts object, allowing users to override or extend default agent instructions. Additionally, the Triage outcome is modeled as a strict sum type (NotABug, Untestable, Testable) with specific fields for each branch, replacing the previous wide wire record to provide clearer, type-safe results for bug reports. A new opaque Title type is also introduced to safely label plan tasks and review issues.
flow/src/main/scala/orca/plan · high confidence
Runner restructured with concurrency guards, per-run trace logging, and worktree support
The runner module has been reorganized into dedicated components to improve reliability and observability. A new concurrency guard (FlowLock) prevents nested or parallel flows from corrupting the git tree by using process-wide flags and workdir lock files. Per-run execution traces are now captured in a rolling log file (OrcaLog) via SLF4J/Logback, with a startup banner (OrcaBanner) displaying the version and log path. Git preconditions are validated upfront (GitPreconditions), and runs can now execute in isolated git worktrees (WorktreeRun) with automatic branch handoff logic (BranchHandoff) to manage HEAD state. Agent wiring (WiredAgents) and flow context (DefaultFlowContext) are centralized, and a LoggingListener ensures all events are mirrored to the trace file.
runner/src/main/scala/orca/runner · high confidence
Test coverage
Added CLI argument parsing and view highlighting tests; Added comprehensive test coverage for the Claude backend; Added comprehensive test coverage for the review and fix loop subsystem; Added comprehensive test coverage for the runner module; Added smoke-test target crates for Orca examples; Added test coverage for agent input serialization, config home resolution, and directory management; Added test coverage for built-in flow compilation, cataloging, and editor integration; Added test coverage for flow orchestration, cost tracking, and diff bounding; Added test coverage for the Gemini CLI backend integration; Added test coverage for the Pi backend integration; Added tests for JSONL event parsing in the Gemini backend; Added tests for Orca Shell UI and Wizard components; Added tests for Orca Shell actions and UI components; Added tests for PR body generation and PR opening flows; Added tests for session manifest reading, round-trip durability, and resume command logic; Added tests for the Claude stream-JSON wire protocol parsers; Added tests for the OpenCode LLM tool backend; Added tests for the Orca Shell flow authoring and sandboxing features; Added tests for the Orca progress tracking subsystem; Added tests for the autonomous planning grid and plan rendering.
Dependencies
Initial build infrastructure and dependency definitions
Establishes the project's build configuration by defining the core dependency versions in \project/Dependencies.scala\ (including Scala 3.9.0, Tapir 1.13.31, Ox 1.0.5, and chimp 0.5.2), setting the sbt version to 1.12.11, and adding necessary plugins such as sbt-softwaremill and sbt-scalafmt. It also introduces a custom \UpdateScalaCliVersionInDocs\ script to automate version bumps in documentation during releases.
project · high confidence
Initial multi-module Scala build and example project scaffolding
The project is bootstrapped as a multi-module Scala build using sbt, introducing distinct modules for shared tools, LLM backends (Claude, Codex, OpenCode, Pi, Gemini), flow execution, and the main runner. This structure establishes the core dependency graph, including libraries like Tapir, Ox, and Chimp, and configures publishing to the org.virtuslab organization. Additionally, example runnable projects are provided, including a basic Rust calculator project defined by a new Cargo.toml.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 68.
Lenses
- Code Health 97
- Architecture 100
- Maturity 85
- Readiness 58
- Security 64
Changes since last survey
- 300 commits — 264 feature/other, 36 fixes
By area
- flow/src — 82 commits
- tools/src — 57 commits
- (root) — 45 commits
- runner/src — 28 commits
- pi/src — 15 commits
- shell/src — 14 commits
- docs/research — 12 commits
- claude/src — 8 commits
- (repo) — 5 commits
- opencode/src — 4 commits
- codex/src — 3 commits
- plans/issue-pr.sc — 3 commits
- adr/0021-orca-shell.md — 2 commits
- docs/plans — 2 commits
- examples/03-bugfix — 2 commits
- examples/issue-pr-bugfix.sc — 2 commits
- examples/runnable — 2 commits
- gemini/src — 2 commits
- plans/issue-pr-bugfix.sc — 2 commits
- .orca/progress-1f5a5271e551.json — 1 commit
Notable commits
- fix: Add file-backed regression test for the gitignored-plan NothingToCommit case
- fix: Add git.diffSince; fix summarisePr's empty-diff bug in both flows
- fix: Add in-memory Plan.implementTaskLoop overload; use it in issue-pr-bugfix.sc
- fix: Clarify issue-pr-bugfix flow by extracting named phases (#5)
- fix: Disable agent auto-commit by default; refresh the PR after the bugfix lands
- fix: Drop sbt hardcodes from issue-pr-bugfix.sc; state both issue flows' outcomes (#29)
- fix: Fix CI race: empty check list is Pending, not Success
- fix: Fix LLM reviewer selection collapsing to zero reviewers
- fix: Fix docs claims about --verbose and the tool-result glyph (#143)
- fix: Fix misleading "lives alongside" docstring in 4 plan scripts
- fix: Fix review loop silently discarding sub-threshold findings (#38)
- fix: Fix the MCP output cap, the unbounded git_show read, rev validation, and two stripMargin sites (#104)
- fix: Fix the agent-isolation findings (#107)
- fix: Fix the cost pipeline (#106)
- fix: Fix the manifest write gate, and stop manifest-less runs evicting continuable ones (#102)
- fix: Measure the fixed per-session preamble and attribute it (T6.2) (#68)
- fix: Merge pull request #1 from zikolach/fix/terminal-control-sanitization
- fix: Migrate example 03 seed to sbt/Scala for issue-pr-bugfix.sc
- fix: Name the failing test for the fix sessions, and thin the flow scripts (#178)
- fix: Reject writes after closeStdin in the piped-process fake, and fix what that exposed (#74)
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
VirtusLab/orca was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit edebe599e084dc65ce847f4bc045740cf331dd9f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.