comet-ml/opik
53.1
Weak · 19 September 2026
595.1k
lines of production code
TypeScript
with Python, Java
1
measurement over time
What this system is
Opik is an open-source platform for observing, evaluating, and optimizing large language model applications. It provides comprehensive distributed tracing and cost intelligence to monitor agent performance, alongside robust evaluation frameworks that support automated benchmarking, LLM-as-a-judge scoring, and human-in-the-loop annotation workflows. The system enables continuous improvement through prompt and agent optimization algorithms, while offering configurable guardrails for security and data privacy.
How it got here
2024 — Opik backend initialization and core feature development
63 changes.
This period established the foundational Opik backend service, introducing the core application infrastructure, database schemas, and essential API endpoints for traces, spans, and datasets. It simultaneously delivered key platform capabilities including Agent Insights, Annotation Queues, Alerting, and LLM cost tracking, while standardizing the development environment and deployment configurations.
2025 — LLM infrastructure expansion and automation features
70 changes.
This period focused on building a comprehensive LLM infrastructure layer, adding support for new providers like OpenRouter and Vertex AI, and enabling multimodal inputs including video and audio. It also introduced significant automation capabilities, such as first-class trace threads, background jobs for insights and alerts, and a new guardrails service for content validation.
2026 — Opik v2 documentation and backend scaling
62 changes.
This period focused on launching the comprehensive Opik v2 documentation site and expanding the platform's evaluation and optimization capabilities with new benchmark packages and algorithm guides. Concurrently, the backend underwent significant scaling work, introducing pre-computed experiment aggregation tables, data retention policies, and local runner integrations, alongside critical database migrations for traces and spans to support high-volume production workloads.
Features
Add OpenTelemetry instrumentation utilities for async operations
The backend now includes a new utility class, InstrumentAsyncUtils, which provides methods to create and manage OpenTelemetry spans for asynchronous operations. This enables distributed tracing for specific product operations by wrapping async code segments with custom spans, allowing users to monitor performance and trace execution flows within the application.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/instrumentation · high confidence
Database schema evolution for core Opik features
This update introduces a comprehensive Liquibase migration suite (migrations 000001–000050) that establishes and evolves the backend database schema to support key platform capabilities. It adds tables and columns for prompt versioning (including template types and tags), LLM provider configuration (supporting custom providers and API keys), and dataset versioning with Git-like tags and evaluation suite support. The schema also enables online scoring automation rules, project-scoped alerts with webhook integrations, workspace dashboards, and trace thread tracking, while adding visibility flags to projects and datasets to control access.
apps/opik-backend/src/main/resources/liquibase/db-app-state · high confidence
Documentation for LLM Guardrails feature
Added comprehensive documentation for the new LLM Guardrails capability, covering the five guard types (PII, Topic, Prompt Injection, LLM Judge, and Custom Classifier), how to define and use guardrail policies in the UI and SDK, and how to fine-tune custom models for specific checks. The new pages also detail the self-hosted Guardrails server configuration, including installation, environment variables, and API usage.
apps/opik-documentation/documentation/fern/docs-v2/guardrails · high confidence
Dynamic token authentication and request interception for custom LLM providers
The backend now supports dynamic OAuth2 client-credentials token fetching for custom LLM providers. A new AuthTokenProvider manages short-lived bearer tokens, caching them in Redis with distributed locking to prevent stampedes, and automatically invalidates cached tokens on 401/403 gateway rejections. The Custom LLM client generator and HTTP client interceptor now apply these dynamic tokens, while also supporting static auth header overrides, URL query parameter injection, and {model} placeholder substitution in base URLs. Error handling is improved with specific mapping for OpenAI-compatible error bodies, including Azure-specific codes.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/customllm · high confidence
Initial Opik backend application setup
The Opik backend application is introduced, establishing the core server infrastructure using the Dropwizard framework. This entry point configures the application lifecycle, including database migrations via Liquibase, dependency injection through Guice, and JSON serialization settings. It integrates essential modules for authentication, caching, rate limiting, and various LLM provider connections (such as OpenAI, Anthropic, and Gemini), while also registering custom exception mappers and parameter converters to handle API requests and errors.
apps/opik-backend/src/main/java/com/comet/opik · high confidence
Initial Python SDK documentation site
The Python SDK documentation site is now available, providing a complete reference for the \opik\ package. It includes guides for installation, configuring the SDK via CLI or Python, and using core features like the \track\ decorator for logging traces and the \evaluate\ function for running evaluations. The reference section documents the main \Opik\ class, context managers, and a dedicated integrations section covering supported frameworks such as LangChain, OpenAI, Anthropic, AWS Bedrock, DSPy, CrewAI, and Guardrails.
apps/opik-documentation/python-sdk-docs/source · high confidence
Initial backend resource configuration and model definitions
This change introduces the foundational resource files for the Opik backend, including a startup banner, an OpenAPI template for API documentation, and Quartz scheduler properties to configure daemon threads and job store behavior. It also adds the default LLM model registry (\llm-models-default.yaml\) and comprehensive pricing data (\model\_prices\_and\_context\_window.json\, \model\_prices\_overrides.json\) for tracking costs across providers like OpenAI, Anthropic, and Gemini.
apps/opik-backend/src/main/resources · high confidence
Initialize analytics database schema with Liquibase migrations
The backend now includes a new Liquibase changelog at \apps/opik-backend/src/main/resources/liquibase/db-app-analytics\ that manages the lifecycle of the analytics database. This change introduces a comprehensive set of SQL migrations that create and configure core ClickHouse tables—including \spans\, \traces\, \feedback\_scores\, \dataset\_items\, \experiments\, \optimizations\, \comments\, \attachments\, \guardrails\, \trace\_threads\, and \annotation\_queues\. The migrations also handle schema evolution tasks such as adding columns (e.g., \model\, \error\_info\, \visibility\_mode\), adjusting data types (e.g., \Decimal\ precision, \DateTime64\ resolution), and migrating tables to \ReplicatedReplacingMergeTree\ engines for cluster support.
apps/opik-backend/src/main/resources/liquibase/db-app-analytics · high confidence
Interactive notebook access for Agent Optimizer cookbooks
The documentation site now injects a banner on specific Agent Optimizer cookbook pages (Quickstart, Hallucination Metric, Moderation Metric, and Dynamic Tracing Control) that provides direct links to view the source on GitHub and run the Jupyter notebooks in Google Colab. This change adds a new client-side script (\add\_cookbook\_banner.js\) to the Fern documentation configuration, enhancing the developer experience by making it easier to experiment with the optimization workflows directly in the browser.
apps/opik-documentation/documentation/fern · high confidence
Introduce Agent Blueprint and Config domain models
The backend now includes the core domain models for the Agent Configuration Management System. This change adds the \AgentConfig\ record to represent project-scoped configuration metadata and the \AgentBlueprint\ record to store blueprint definitions (including type, name, description, and key-value pairs). It also introduces the \AgentConfigDAO\ with data access methods to persist and retrieve these entities from the \agent\_configs\ and \agent\_blueprints\ database tables, laying the groundwork for managing agent blueprints and their associated environments.
apps/opik-backend/src/main/java/com/comet/opik/domain · high confidence
Introduce Evolutionary Optimizer with DEAP-based genetic algorithms
The Opik Optimizer SDK now includes an EvolutionaryOptimizer that uses the DEAP library to evolve and improve prompts over multiple generations. This new capability supports both single-objective and multi-objective optimization (balancing performance and prompt length), featuring genetic operations like crossover, mutation, and selection. It also introduces support for multimodal inputs and tool optimization within the evolutionary process, providing a new algorithmic approach for prompt engineering alongside existing optimizers like GEPA and MetaPrompt.
_sdks/opik\optimizer · high confidence
Introduce Local Runner API models and prompt mask support
The backend now exposes a new set of API models for the Local Runner integration, defining the data structures for runner registration, job submission, and command execution. This includes new request and response records for managing local runner jobs (e.g., \CreateLocalRunnerJobRequest\, \LocalRunnerJob\) and a bridge command system (\BridgeCommand\, \BridgeCommandType\) that allows agents to perform file and execution operations. A key behavioral addition is the support for non-destructive prompt resolution: the job models now accept a \promptMasks\ map to apply mask overlays keyed by prompt ID, while the legacy \maskId\ field is marked as deprecated. The entry also introduces enums for runner status, job status, and command types to standardize the interaction between the backend and local runner agents.
apps/opik-backend/src/main/java/com/comet/opik/api/runner · high confidence
Introduce Opik Guardrails Backend service with PII, topic, and prompt injection validation
A new backend service for content guardrails is introduced, providing a Flask API at /api/v1/guardrails/validations that accepts text and runs multiple validation types in a single request. The service supports PII detection (using Presidio), topic classification (using BART), and prompt injection detection (using a fine-tuned Qwen classifier). It also includes endpoints for training and serving custom binary classifiers via LoRA adapters. The service is containerized with two Dockerfiles: a GPU-enabled image (CUDA 12.2) and a CPU-only image (Python slim), both running as a non-root user (UID 1001) and pre-caching required models. Configuration is driven by environment variables such as OPIK\_GUARDRAILS\_DEVICE and OPIK\_GUARDRAILS\_PROMPT\_INJECTION\_BASE\_MODEL.
apps/opik-guardrails-backend · high confidence
Introduce Opik Python Backend service for sandboxed code execution and optimization jobs
This change introduces the \opik-python-backend\ service, a new component responsible for running Python code in isolated environments and executing Optimization Studio jobs. The service provides a Docker-based execution strategy (via \DockerExecutor\) and a subprocess-based strategy (via \IsolatedSubprocessExecutor\), allowing users to run metrics and evaluators in secure, isolated contexts. It includes infrastructure for the Optimization Studio, consuming jobs from a Redis queue, running optimization algorithms in separate subprocesses with isolated environment variables, and streaming logs back to the UI. The service also handles demo data generation and seeding, ensuring consistent initial state for new workspaces. Configuration is managed via environment variables for execution strategy, parallelism, timeouts, and resource limits.
apps/opik-python-backend · high confidence
Introduce Opik pytest plugin for LLM unit testing
Adds a new pytest plugin that enables tracking LLM-based tests as experiments. Users can mark test functions with the \@llm\_unit\ decorator to automatically capture test inputs, expected outputs, and metadata as traces. Upon test completion, the plugin logs a 'Passed' feedback score to each trace and aggregates the results into an Opik experiment (stored in a 'tests' dataset), providing a summary in the terminal with a link to the Opik UI.
sdks/python/src/opik/plugins/pytest · high confidence
Introduce Redis-based distributed locking, rate limiting, and caching infrastructure
The backend now uses Redis for core distributed coordination and caching capabilities. A new LockService provides distributed locks and semaphore-based concurrency control with OpenTelemetry metrics (lock\_waiting, lock\_held, lock\_acquire\_wait\_milliseconds) to monitor contention and prevent deadlocks. A RedisRateLimitService implements rate limiting using Redisson's rate limiters, supporting configurable limits and durations. A RedisCacheManager provides a unified caching layer with both reactive and non-reactive APIs for storing and retrieving cached data with TTLs. Additionally, a LenientUUIDDeserializer ensures backward compatibility with Redis stream messages by tolerating both plain-string and Jackson polymorphic UUID formats, preventing stream processing failures during upgrades.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/redis · high confidence
Introduce anonymous usage reporting and product analytics infrastructure
The backend now collects anonymous usage data and product analytics events to help improve the product. This includes a daily usage report sent at midnight (tracking total/daily users, traces, experiments, and datasets), an installation/startup event sent once per instance, and a 'first trace created' event per workspace. Users can disable these reports via the existing usage-report and analytics configuration flags; when disabled, no telemetry is sent. The implementation uses an async, non-blocking HTTP client to ensure analytics failures do not impact API performance.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/bi · high confidence
Introduce backend cost tracking for LLM model usage
The backend now calculates and reports costs for LLM spans and traces based on token usage. This change introduces a new CostService that loads pricing data from JSON files and applies provider-specific logic to handle input, output, audio, reasoning, and cached tokens. It supports a wide range of providers including OpenAI, Anthropic, Bedrock, Google AI, Azure, and others, with specific handling for tiered pricing (e.g., above 128k/200k/272k tokens) and cache discounts. Model name normalization is applied to handle variations like dot-separated names, date suffixes, and provider prefixes, ensuring accurate price lookups. The system also falls back to metadata-based costs if calculation fails or returns zero.
apps/opik-backend/src/main/java/com/comet/opik/domain/cost · high confidence
Introduce data access and mapping layers for annotation queue automation rules
Added new backend components to support the annotation queue router feature within the automation rules system. This includes DAOs for persisting and querying router configurations (including joins to the parent automation rules table), a MapStruct mapper to convert between storage models and API objects, and a model record representing the router entity. These files establish the storage and data-binding infrastructure for routing annotations to queues based on configurable conditions and scopes.
apps/opik-backend/src/main/java/com/comet/opik/domain/evaluators · high confidence
Introduce data retention policy configuration
The backend now supports defining data retention policies through new API models. Users can configure retention rules at the organization, workspace, or project level using predefined periods (14 days, 60 days, 400 days, or unlimited). The system allows enabling these rules to apply to past data and tracks the progress of historical cleanup operations via catch-up cursors.
apps/opik-backend/src/main/java/com/comet/opik/api/retention · high confidence
Introduce first-class trace threads with status tracking and lifecycle management
Threads are now first-class entities in the system, allowing users to track their status (active, inactive, closed) and manage their lifecycle. This change introduces new backend components (TraceThreadService, TraceThreadDAO, TraceThreadIdService) to handle thread creation, closure, and status updates, including support for online scoring and bulk operations. Users can now explicitly open and close threads, and the system automatically marks threads as inactive based on configurable timeouts. The implementation also includes optimizations for query performance and consistency when handling thread data.
apps/opik-backend/src/main/java/com/comet/opik/domain/threads · high confidence
Introduce lightweight metric base package for fast imports
The Python SDK now includes a lightweight \\_opik\ package that exposes \BaseMetric\ and \ScoreResult\ without loading the full Opik client or heavy dependencies. This allows users to import and use metric types with near-zero import time, which is particularly beneficial for environments where fast startup is critical or where only metric definitions are needed without full tracking overhead.
sdks/python · high confidence
Introduce pre-computed experiment aggregation tables for faster experiment views
The backend now supports pre-computed aggregation tables for experiments, allowing experiment lists, groups, and dataset item comparisons to be served from denormalized metrics rather than joining raw trace and span data on every request. This change introduces a new aggregation pipeline (including an aggregation publisher and background job infrastructure) that computes and stores metrics such as trace durations, span costs, feedback scores, and pass rates in ClickHouse, significantly reducing query latency and database load for experiment-related endpoints.
apps/opik-backend/src/main/java/com/comet/opik/domain/experiments · high confidence
Introduce user-facing logging to ClickHouse
The backend now captures user-facing application logs and stores them in ClickHouse. This is implemented via a new \ClickHouseAppender\ that buffers log events and flushes them asynchronously to the database, a \UserFacingLoggingFactory\ to configure the logging pipeline, and a \LogContextAware\ utility to ensure MDC context is preserved across thread boundaries in reactive streams.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/log · high confidence
Introduces a new structured filter API for all domain entities
The backend now exposes a unified, structured filter API across all major entities (Alerts, Annotation Queues, Automations, Dashboards, Datasets, Experiments, Optimizations, Prompts, Spans, and Traces). Instead of ad-hoc query parameters, users can now filter resources using a consistent JSON structure defining the field, operator, and value. This change introduces support for a wide range of filterable fields per entity (e.g., filtering Spans by cost, duration, or feedback scores; filtering Datasets by tags or type) and includes validation logic to ensure data integrity for types like dates, numbers, and enums.
apps/opik-backend/src/main/java/com/comet/opik/api/filter · high confidence
Introduces application event infrastructure with automatic listener registration and observability
The backend now includes a new event-driven architecture in the infrastructure layer, featuring an asynchronous event bus that automatically discovers and registers listeners annotated with @Subscribe. This change adds built-in observability for event processing, tracking listener invocations and errors via OpenTelemetry metrics (opik.event.invoked, opik.event.error) and distributed tracing, while providing a base event class that automatically captures trace context and workspace metadata for all published events.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/events · high confidence
Introduces configurable rate limiting for API endpoints
The backend now enforces rate limits on API requests using a new infrastructure layer. This change adds a \RateLimited\ annotation that developers can apply to methods to enable throttling, supporting limits scoped to the workspace, the user (via API key or CIPX device ID), and custom buckets. The system integrates with OpenTelemetry to expose metrics for rate-limited requests and ensures standard rate-limit headers (including the reset time) are correctly propagated in HTTP responses.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/ratelimit · high confidence
Introduces debounced, project-scoped alerting with webhook delivery and event logging
The backend now supports a new alerting mechanism that aggregates alert events into debounced buckets stored in Redis, preventing notification spam by grouping events by alert ID and type within a configurable time window. Alerts can be scoped to specific projects (using a new project\_id column or legacy trigger configs) and, for guardrails, filtered by specific guardrail types. When a bucket is ready, a consolidated webhook notification is sent containing aggregated event details, and all alert events are persisted to an alert\_logs table for auditability and retrieval.
apps/opik-backend/src/main/java/com/comet/opik/domain/alerts · high confidence
Introduces factory for user-facing log table data access
The backend now includes a factory interface and implementation that maps specific user log types (automation rule evaluator and alert event) to their respective database access objects. This infrastructure change enables the system to persist user-facing logs into the underlying storage layer, supporting the new logging capabilities for automation rules and alerts.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/log/tables · high confidence
Introduces infrastructure configuration classes for backend services
The backend now exposes a comprehensive set of configuration classes in the infrastructure package, allowing operators to tune various system behaviors. These include settings for agent insights reports, analytics, authentication timeouts and retries, batch operations, caching, CORS, ClickHouse database connectivity (including async insert tuning and distributed table wrapping), dataset export and versioning, distributed locking, encryption, and experiment aggregation. This change enables fine-grained control over performance, reliability, and feature toggles without code changes.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure · high confidence
Introduction of attachment data models and API contracts
The backend now supports attaching files to traces and spans. This change introduces the core API data structures required for the feature, including the \Attachment\ record for paginated lists, \AttachmentInfo\ for metadata, and specific request/response models for the multipart upload workflow (\StartMultipartUploadRequest\, \StartMultipartUploadResponse\, \CompleteMultipartUploadRequest\, \MultipartUploadPart\) as well as deletion (\DeleteAttachmentsRequest\). It also defines the \EntityType\ enum to distinguish between trace and span attachments.
apps/opik-backend/src/main/java/com/comet/opik/api/attachment · high confidence
MCP OAuth 2.1 Authorization Server endpoints and configuration
The backend now exposes a configurable OAuth 2.1 authorization server for Model Context Protocol (MCP) clients, enabled via the \mcpOAuth.enabled\ configuration flag. This adds several new REST endpoints: a metadata endpoint at \/.well-known/oauth-authorization-server\ (RFC 8414), a dynamic client registration endpoint at \/oauth/register\ (RFC 7591) with IP-based rate limiting, an authorization endpoint at \/oauth/authorize\ that handles PKCE validation and redirects to login or consent screens, a consent context endpoint at \/oauth/authorize/context\, and a token endpoint at \/oauth/token\ supporting authorization code and refresh token flows. It also includes a token revocation endpoint at \/oauth/revoke\ (RFC 7009) and a bearer token validation endpoint at \/oauth/validate\. The feature is wired into the application via \McpOAuthBundle\, which conditionally registers these resources based on the configuration.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/oauth · high confidence
MCP OAuth dynamic client registration and connection tracking
The backend now supports RFC 7591 dynamic client registration for MCP hosts, allowing them to register their metadata (such as software ID and logo) and redirect URIs. This change introduces a new connection tracking system that records which hosts are connected to a workspace, persists their display metadata, and correctly identifies new connections versus re-authorizations. It also implements secure token handling with PKCE, family-based refresh token rotation to handle concurrent requests, and a background scrub job to clean up expired codes and tokens.
apps/opik-backend/src/main/java/com/comet/opik/domain/mcpoauth · high confidence
Native Slack and PagerDuty alert integrations
Users can now receive alerts directly in Slack and PagerDuty. The backend introduces dedicated payload mappers and data models for both services: Slack notifications are formatted using Block Kit with structured headers, summaries, and detailed metric or event information, while PagerDuty notifications adhere to the Events API v2 structure with appropriate severity levels. This change adds the necessary backend components to serialize and send webhook events to these external platforms.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/webhooks/slack · high confidence
New /is-alive and /ver health check endpoints
The backend now exposes a new /is-alive endpoint that reports server liveness based on critical health checks, ensuring that non-critical failures (such as optional ClickHouse connectivity issues) do not incorrectly mark the server as down for SDKs and frontend clients. Additionally, a /ver endpoint is available to retrieve the current application version.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/health · high confidence
New API data models for Agent Insights, Annotation Queues, and Alerts
The backend API layer introduces a set of new request and response models to support Agent Insights, Annotation Queues, and Alerting. For Agent Insights, records like \AgentInsightsJob\, \AgentInsightsIssue\, and \AgentInsightsReport\ define the data structures for managing automated agent scans, tracking discovered issues with severity and status, and ingesting daily report results. Annotation Queue support is added via \AnnotationQueue\ and \AnnotationQueueAutomation\, enabling the definition of queues with configurable scopes, automation rules based on feedback scores, and item limits. Finally, the Alerting system is modeled with \Alert\, \AlertTrigger\, and \AlertTriggerConfig\, supporting various event types (such as trace errors, cost, and latency) and configuration types (like thresholds and project scoping) to allow users to set up metric-based notifications.
apps/opik-backend/src/main/java/com/comet/opik/api · high confidence
New Agent Config and Annotation Queue APIs in TypeScript SDK
The TypeScript SDK now includes support for Agent Configurations and Annotation Queues. Users can manage versioned agent configurations (blueprints and masks) via the new ConfigManager, Blueprint, and Config classes, which handle value resolution, prompt linking, and automatic metadata injection into traces. Additionally, the SDK introduces Annotation Queue support, allowing users to add, remove, and retrieve traces or threads for review using the TracesAnnotationQueue and ThreadsAnnotationQueue classes.
sdks/typescript · high confidence
New Comet plugin components for assistant, billing, and workspace management
The \apps/opik-frontend/src/plugins/comet\ directory now includes a suite of new UI components and providers that power the assistant sidebar, workspace collaboration, and billing workflows. This includes the \AssistantSidebar\ and its supporting states (\AssistantErrorState\, \OllieLoader\, \AssistantDebugInfo\) for loading, error handling, and version display, as well as the \CollaboratorsTab\ and \InviteUsersPopover\ for managing workspace members and permissions. Additionally, \BillingLink\ and \RetentionBanner\ provide admin-only access to credits and usage warnings, while \LayoutProvider\ and \PermissionsProvider\ supply context for AI spend and granular feature permissions to the rest of the application.
apps/opik-frontend/src/plugins/comet · high confidence
New Cost Intelligence documentation covering installation, data privacy, and policy management
Added a comprehensive documentation section for the Cost Intelligence feature, detailing how to install the plugin via MDM, managed settings, or the macOS app, explaining the data privacy model (metadata-only collection with optional content capture), and describing how organization and user cost policies are rolled out and enforced to reduce coding agent spend.
_apps/opik-documentation/documentation/fern/docs-v2/cost\intelligence · high confidence
New Cursor extension for Opik integration
The Cursor extension for Opik is now available, allowing you to automatically save Cursor chat sessions to Opik for team sharing and analysis. New Cursor chats are uploaded automatically, while historical chats can be imported manually via the command palette. The extension requires an Opik API key for configuration and retains request identities for 180 days to prevent duplicate uploads.
extensions · high confidence
New Docker Compose deployment with service profiles and local development overrides
The deployment directory now includes a complete \docker-compose.yaml\ stack and supporting configuration files, introducing Docker Compose profiles to manage service combinations (infrastructure, backend, full Opik suite, guardrails, and OpenTelemetry). It adds specific override files (\docker-compose.override.yaml\, \docker-compose.local-be.yaml\, \docker-compose.local-be-fe.yaml\) to expose internal service ports for local debugging and to support running the backend or frontend locally while using Docker for the rest of the stack. The configuration also includes ClickHouse settings to align the local environment with Kubernetes deployments (e.g., distributed insert queues) and Nginx configurations to handle routing for the main application, guardrails, and OAuth flows.
deployment · high confidence
New LLM infrastructure layer with model registry and Gemini thinking support
The backend introduces a new LLM infrastructure package that adds a configurable LLM model registry (loaded from classpath, remote YAML, or local override) with a scheduled refresh job, and a factory-based provider routing system that resolves providers via the registry before falling back to enum-based matching. It adds native support for Gemini thinking parameters (level vs. budget mapping for Gemini 3+), an Anthropic client configuration record, and a delegating HTTP client builder for timeout forwarding. Error handling is standardized with OpenAI-compatible status code mapping and structured output strategy selection (tool calling vs. instruction) driven by the registry. Logging is improved via a streaming response logger that assembles and logs complete responses and token usage, and the LangChain mapper is updated to support multimodal content including video MIME type detection.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm · high confidence
New OpenTelemetry mapping infrastructure for AI integrations
The backend now includes a dedicated mapping layer for OpenTelemetry data, introducing classes to process span events into metadata, define structured mapping rules (supporting prefix and exact matches), and aggregate rules from various AI integrations like LiveKit, GenAI, and Claude Code. This infrastructure enables more accurate extraction of usage, cost, and model details from OTel traces, including handling pre-computed costs and normalizing token counts for cache pricing.
apps/opik-backend/src/main/java/com/comet/opik/domain/mapping · high confidence
New Opik 2.0 E2E test infrastructure and coverage taxonomy
The \tests\_end\_to\_end\ directory now includes the foundation for the Opik 2.0 E2E test suite, featuring a new capability taxonomy (\taxonomy.yaml\) and a tag-based coverage system (\TESTING-TAGS.md\). This system uses Playwright tags (\@area\, \@cap\, \@vcap\) to map tests to product capabilities, with automated linting (\tag\_lint.py\) and reconciliation (\reconcile.py\) to maintain coverage accuracy. The suite supports multiple deployment environments (OSS, Cloud, Self-hosted) via environment configuration and includes fixtures for local agent runners and Ollie integration testing.
_tests\_end\_to\end · high confidence
New Opik Python SDK E2E driver service
A new FastAPI-based bridge service has been added to the e2e test infrastructure to drive the Opik Python SDK programmatically. This service exposes REST endpoints for creating and managing core resources—projects, datasets, traces (including nested spans and backdated timestamps), prompts, test suites, annotation queues, and feedback definitions—and provides specialized evaluation routes for experiments and conversation threads. It handles SDK client lifecycle, authentication, and specific test requirements such as JSON request compression toggling and deterministic scoring, enabling comprehensive end-to-end smoke and regression tests for these SDK capabilities.
_tests\_end\_to\end/e2e/services/opik-sdk-driver · high confidence
New Opik v2 documentation site adds comprehensive contributor guides
The Opik documentation site has been updated to version 2, introducing a new set of contributor guides located in the \contributing\ section. These new pages provide detailed instructions for setting up local development environments (including Docker, local process, and manual modes) for the Java backend, React frontend, Python SDK, and TypeScript SDK. The update also adds specific guides for contributing to the Agent Optimizer SDK, general documentation standards, and information on the Opik Bounty Program.
apps/opik-documentation/documentation/fern/docs-v2/contributing · high confidence
New Prompt Library and MCP server documentation
The documentation now includes dedicated guides for the Prompt Library and the Opik MCP server. Users can learn how to manage, version, and fetch prompts using the Python and TypeScript SDKs, as well as how to use the Prompt Playground to test variants. Additionally, a new guide details the one-command setup (\uvx opik mcp configure\) for connecting AI coding assistants like Cursor and VS Code to Opik for automated instrumentation and trace analysis.
_apps/opik-documentation/documentation/fern/docs-v2/prompt\engineering · high confidence
New Python SDK load-test suite and cutover rehearsal tooling
The \tests\_load\ directory now provides a structured pytest-driven load-test suite for the Python SDK and CLI-based rehearsal tools for backend data-cutover operations. The SDK suite (\suite/python\_sdk/\) exercises ingestion performance across eight scenarios: high-rate trace/span ingestion, heavy payloads (up to 1 MB per trace/span), explicit and implicit (base64-extracted) attachments, burst and spread traffic patterns, concurrent multi-threaded writers sharing a single client, and sequential dataset version inserts. Each test logs metrics to JSON, supports scaling via \--load-scale\, and verifies that all submitted items are retrievable. Additionally, the \tests/cutover\_common/\ package supplies shared helpers and a delete-traffic generator to rehearse traces and spans cutover runbooks locally, including deletion replay and resurrection guard exercises against a local ClickHouse instance.
_tests\load · high confidence
New RQ queue infrastructure for Python integration
Added a new Redis Queue (RQ) infrastructure layer in the backend to support Python integration jobs. This includes data models for job definitions and statuses, a producer interface for enqueuing jobs, and utilities to build RQ-compatible Redis HASH structures. A key behavioral detail is the timestamp formatting: the new mapper truncates Java Instant timestamps to microsecond precision (6 decimal places) to ensure compatibility with Python RQ's parsing expectations, preventing serialization errors that would occur with Java's default nanosecond precision.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/queues · high confidence
New TypeScript SDK evaluation documentation and API reference
The TypeScript SDK reference documentation has been expanded to cover the full evaluation workflow, including creating and managing datasets with versioning, running evaluations via the \evaluate\ and \evaluatePrompt\ functions, configuring metrics and models, and managing experiments. A new section details the Opik Query Language (OQL) for filtering traces and prompts, and the \opik-ts\ CLI tool is documented for automated project setup. These changes provide users with comprehensive guides and API references for building, testing, and monitoring LLM applications using the Opik TypeScript SDK.
apps/opik-documentation/documentation/fern/docs-v2/reference/typescript-sdk · high confidence
New advanced evaluation documentation pages added
The documentation site now includes five new advanced evaluation guides: Annotation Queues (managing SME review workflows), Agent Trajectory Evaluation (assessing tool selection and reasoning paths), Multi-Turn Agent Evaluation (using LLM-based user simulation for chatbots), Experiment Result Export (SDK-based CSV export for large datasets), and Resuming Interrupted Evaluations (recovering from crashes or network issues).
apps/opik-documentation/documentation/fern/docs-v2/evaluation/advanced · high confidence
New advanced optimization documentation for Opik v2
The Opik v2 documentation now includes a new 'advanced' section for optimization runs, providing detailed guides on chaining optimizers, extending the framework with custom algorithms, and configuring custom metrics. It also introduces documentation for new v3.0.0+ sampling controls, including dataset subsampling strategies and the 'n' parameter for generating multiple completions per evaluation to reduce variance.
apps/opik-documentation/documentation/fern/docs-v2/development/optimization-runs/advanced, apps/opik-documentation/documentation/fern/docs-v2/development/optimization-runs/cookbooks, apps/opik-documentation/documentation/fern/docs-v2/reference/python-sdk · high confidence
New advanced tracing documentation pages added
The documentation site now includes a comprehensive set of new advanced tracing guides covering user feedback logging, cost tracking, agent graph visualization, chat conversation threading, distributed tracing (including OpenTelemetry bridging), multimodal media attachments, offline fallback with message replay, and detailed SDK configuration for both Python and TypeScript.
apps/opik-documentation/documentation/fern/docs-v2/tracing · high confidence
New annotation-based caching infrastructure with deadlock prevention
The backend now includes a new caching infrastructure in the \infrastructure.cache\ package, introducing \@Cacheable\, \@CachePut\, and \@CacheEvict\ annotations to declaratively manage cache interactions. This system is backed by a \CacheInterceptor\ and \CacheManager\ that support both reactive (Mono/Flux) and synchronous execution paths. A critical behavioral safeguard is built into these annotations: they explicitly warn against and prevent nested cache operations (e.g., annotating both a delegating method and its implementation), which previously caused deadlocks, Reactor threading violations, and Redis timeouts. This change provides the foundational mechanism for future caching features, such as those referenced in automation rule evaluation and analytics deduplication.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/cache · high confidence
New attachment storage and processing infrastructure
The backend now includes a dedicated attachment domain to handle file storage and processing. This introduces services to detect and strip base64-encoded attachments from trace and span payloads (AttachmentStripperService) and reinject them when requested (AttachmentReinjectorService). It adds data access and business logic layers (AttachmentDAO, AttachmentService) to manage attachment metadata in the database and handle uploads/downloads via S3/MinIO using multipart uploads and presigned URLs (FileService, PreSignerService).
apps/opik-backend/src/main/java/com/comet/opik/domain/attachment · high confidence
New backend utility library for async context, JSON handling, and data formatting
The backend now includes a new set of utility classes in the \com.comet.opik.utils\ package to standardize common operations. \AsyncUtils\ manages Reactor context propagation for user and workspace identity, while \JsonUtils\ centralizes Jackson \ObjectMapper\ configuration, including stream-read constraints and custom deserializers for OpenAI messages and durations. \SentinelTranslation\ handles the mapping between Java nulls and ClickHouse sentinel values (epoch, NaN, empty strings) for non-nullable analytics columns. Additional utilities include \TruncationUtils\ for JSON size limits, \ChunkedOutputHandlers\ for SSE streaming, \RetryUtils\ for transient error handling, and \FileNameUtils\ for safe Content-Disposition header generation.
apps/opik-backend/src/main/java/com/comet/opik/utils · high confidence
New background jobs for Agent Insights, metrics alerts, and system maintenance
The backend now includes several new scheduled jobs to handle automated reporting, alerting, and maintenance tasks. The AgentInsightsReportJob runs daily to generate and publish Agent Insights reports for projects with recent traces, using distributed locking to ensure idempotency across replicas. The MetricsAlertJob evaluates cost and latency thresholds for configured alerts, firing webhooks when thresholds are breached while preventing duplicate notifications via per-alert locking. The AlertJob processes debounced alert buckets from Redis every 5 seconds to trigger consolidated webhook notifications. Additionally, the system introduces jobs for infrastructure maintenance and observability: ClickHousePartitionMetricsJob publishes partition-health metrics to OpenTelemetry, ExperimentDenormalizationJob flushes debounced experiment aggregation events to a Redis stream, DatasetVersionItemsTotalMigrationJob performs a one-time migration for dataset version item counts, and LocalRunnerReaperJob cleans up dead local runner instances. The OllieDailyReportJob and OllieReportMetricsJob handle daily report triggering and pending report metrics respectively.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/jobs · high confidence
New benchmark packages for HotpotQA, HoVer, IFBench, and PUPA
The Opik Optimizer SDK now includes dedicated benchmark packages for four new evaluation datasets: HotpotQA, HoVer, IFBench, and PUPA. The HotpotQA package provides a multi-hop retrieval agent with BM25-based Wikipedia search and specific exact-match/F1 metrics. The HoVer package adds claim-verification metrics and an LLM-based judge for supported/not-supported classification. The IFBench package introduces a constraint-compliance judge to audit instruction-following tasks. Finally, the PUPA package provides dual metrics for privacy-preserving rewrites, including a quality judge and a PII leakage ratio calculator. These packages integrate with the existing benchmark runner architecture to allow users to evaluate and optimize agents against these specific benchmarks.
(repo-wide) · high confidence
New data model for automation rule evaluators
The backend introduces a new set of API data models for automation rules, defining a sealed hierarchy that supports multiple evaluator types (LLM-as-Judge, User-Defined Python Metrics) across different scopes (Traces, Spans, Threads). This change adds the core request/response structures for creating and updating these rules, including fields for sampling rates, trigger scopes, and project assignments, laying the groundwork for the online evaluation feature.
apps/opik-backend/src/main/java/com/comet/opik/api/evaluators · high confidence
New database infrastructure components for data persistence and validation
The backend now includes a comprehensive set of new database infrastructure classes in the \infrastructure.db\ package to support recent feature additions. This includes a \UuidV7TimestampValidator\ that enforces ingestion timestamp windows to prevent partition corruption, with support for audit-only modes and workspace-scoped bypasses. A new \DatabaseAnalyticsModule\ wires up a read-only ClickHouse client specifically for Agent Insights freeform SQL queries. Additionally, numerous JDBI mappers and argument factories have been added to handle serialization and deserialization for new data types, including enums (e.g., \DashboardType\, \AlertType\), JSON structures (e.g., \PromptVersion\, \ExecutionPolicy\), and encrypted configurations (\ProviderAuthConfig\, \EncryptedAuthConfig\).
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/db · high confidence
New development scripts and tooling for local Opik workflows
This change introduces a comprehensive suite of new scripts in the \scripts/\ directory to streamline local development and CI processes. The primary addition is the \dev-runner.sh\ (and its PowerShell counterpart \dev-runner.ps1\), a unified development environment runner that manages Docker infrastructure, backend, and frontend services. It supports multiple modes including standard, BE-only, and an opt-in EM/Platform mode (via \PLATFORM\_ENABLED=true\) that integrates the Comet stack. The runner features multi-worktree support with automatic port offsetting to prevent collisions, and includes flags for quick restarts, database migrations, and debugging. Additionally, new scripts provide specialized tooling: \generate\_openapi.sh\ and \start\_openapi\_server.sh\ for OpenAPI spec generation and local testing; \check-public-fe-plugins.sh\ to prevent committing private frontend plugins; \check\_backend\_migration\_conflicts.sh\ and \check\_clickhouse\_migrations\_cluster.sh\ to validate database migration prefixes and ClickHouse DDL syntax; \analyze\_trivy\_report.sh\ for Docker vulnerability analysis; and \convert-frontmatter.sh\/\convert-mcp.sh\ for AI agent configuration interoperability. A \README.md\ documents these scripts and their usage.
scripts · high confidence
New documentation for Opik 2.0 evaluation capabilities
The documentation site now includes a complete set of new pages for the evaluation section, covering the two core evaluation approaches: Test Suites (assertion-based testing) and Datasets & Metrics (quantitative scoring). Users can now find guides on getting started, building test suites via the SDK, UI, or Ollie, evaluating agents, and specific cookbooks for hallucination and moderation metrics. Additionally, a new guide explains how to evaluate conversation threads using the \evaluate\_threads\ function in the Python SDK.
apps/opik-documentation/documentation/fern/docs-v2/evaluation · high confidence
New documentation for Opik Agent Optimizer workflows
Added a comprehensive set of documentation pages for the Opik Agent Optimizer, covering core concepts, LLM provider configuration, dataset and metric definition, and optimization strategies for prompts, agents, multimodal inputs, and MCP tools.
apps/opik-documentation/documentation/fern/docs-v2/development/optimization-runs/optimization · high confidence
New documentation for Opik administration and authentication
The documentation site now includes comprehensive guides for the Administration section, covering the Admin Dashboard overview, user and workspace management, service accounts, and the full roles and permissions model. It also introduces detailed configuration instructions for authentication methods, including SAML SSO, OIDC SSO, and JWT authentication for programmatic access.
apps/opik-documentation/documentation/fern/docs-v2/administration · high confidence
New documentation for Opik evaluation metrics and configuration
The documentation site now includes detailed guides for the Opik evaluation metrics located in the \evaluation/metrics\ section. New pages cover advanced configuration options such as asynchronous scoring (\ascore\), controlling evaluator temperature, log-probability handling, and tracking controls. Additionally, reference documentation has been added for specific metrics including Agent Task Completion, Agent Tool Correctness, Answer Relevance, Compliance Risk, Context Precision, Context Recall, and a comprehensive set of Conversational Metrics (e.g., Knowledge Retention, Session Completeness, User Frustration). The docs also explain how to create custom single-turn and multi-turn conversation metrics, as well as how to integrate custom LLM providers using LiteLLM.
apps/opik-documentation/documentation/fern/docs-v2/evaluation/metrics · high confidence
New documentation for production alerts, anonymizers, and online evaluation rules
The production documentation site now includes dedicated pages for configuring automated webhook alerts (with native Slack and PagerDuty integrations), protecting sensitive data via PII anonymizers, and defining online evaluation rules using LLM-as-a-Judge metrics. These pages provide step-by-step setup guides, code examples for custom rules and anonymizers, and details on variable mapping and agent-based evaluation for production traces.
apps/opik-documentation/documentation/fern/docs-v2/production · high confidence
New documentation for prompt optimization algorithms and benchmarking
Added comprehensive documentation pages for the Opik prompt optimization algorithms, including guides for the Evolutionary, Few-Shot Bayesian, GEPA, HRPO, MetaPrompt, and Parameter optimizers. The new content covers quickstart code snippets, configuration options, model support, and usage notes for each algorithm. Additionally, a new benchmarks page was added to help users compare algorithm performance across common datasets and reproduce results using the public benchmark scripts.
apps/opik-documentation/documentation/fern/docs-v2/development/optimization-runs/algorithms · high confidence
New domain event types for backend infrastructure
The backend now exposes a comprehensive set of new domain event classes in the \com.comet.opik.api.events\ package to support internal asynchronous processing, distributed scoring, and observability. These include events for entity lifecycle changes (traces, spans, comments, experiments, datasets, optimizations), attachment uploads, and cost intelligence updates. Additionally, new event types facilitate LLM-as-Judge and Python-based scoring for traces, spans, and threads, while a new \RedisSubscriberMessage\ interface and \BaseEvent\ base class standardize message handling for the Redis-based subscriber infrastructure.
apps/opik-backend/src/main/java/com/comet/opik/api/events · high confidence
New event-driven infrastructure for Agent Insights, attachment processing, and cost intelligence
This change introduces a suite of new event subscribers and listeners in the backend to decouple and scale key platform capabilities. The Agent Insights Report Subscriber now consumes triggers from a Redis stream to generate daily or manual reports, handling failures gracefully by marking runs as failed without blocking the queue. A new Agentic Scoring Service centralizes the tool-call loop and attachment handling logic previously scattered across scorers, enabling trace, span, and thread-level LLM-as-judge evaluations to share a robust implementation that tolerates attachment upload races. Attachment uploads are now processed asynchronously via an Attachment Upload Listener, which decouples file processing from the main request flow and provides detailed OpenTelemetry metrics. Additionally, a Cost Intelligence Ingestion Listener automatically populates dedicated spend and identity tables (cipx\_spends, cipx\_trace\_identities) as spans and traces are created, while a Closing Trace Thread Subscriber manages the automatic closure of trace threads based on configured timeouts. These components all extend a new Base Redis Subscriber, which provides a standardized, resilient foundation for consuming Redis streams with proper error handling, metrics, and cursor management.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events · high confidence
New integration documentation for Google ADK, AG2, Agent Spec, Agno, AISuite, Anthropic, AutoGen, Bedrock, and BeeAI
The documentation site now includes dedicated integration guides for a broad set of new AI frameworks and providers. Users can now find setup instructions and code examples for Google Agent Development Kit (ADK), AG2, Agent Spec, Agno, AISuite, Anthropic, AutoGen, AWS Bedrock, and BeeAI (TypeScript). These pages cover installation, configuration for Cloud/Enterprise/Self-hosted deployments, and specific instrumentation patterns such as automatic agent tracking, OpenTelemetry setup, and cost tracking.
apps/opik-documentation/documentation/fern/docs-v2/integrations · high confidence
New internal analytics and usage endpoints for system monitoring
Two new internal REST resources have been added to expose system metrics and analytics capabilities. The \/v1/internal/usage\ endpoint now provides daily counts for traces and spans per workspace, a detailed spans breakdown by workspace/project/user, and BI event information for traces, experiments, datasets, and spans. Additionally, the \/v1/internal/analytics-queries\ endpoint allows executing free-form, read-only SQL queries against ClickHouse for Agent Insights, which is gated behind the \ollieEnabled\ service toggle and enforces workspace/project scoping.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/internal · high confidence
New observability documentation pages for v2
The documentation site now includes a comprehensive set of new pages under the observability section, covering the full agent lifecycle: an overview of observability capabilities, a getting-started guide for adding tracing via AI coding agents or manual SDK integration, a detailed guide for the Agent Playground (running and debugging agents locally), instructions for debugging agents with Ollie, a new Diagnostics page for automated issue detection, documentation for the \opik migrate\ CLI for moving data between projects, and an export data guide. These pages provide users with the reference material needed to instrument, monitor, debug, and manage their LLM applications within the Opik v2 platform.
apps/opik-documentation/documentation/fern/docs-v2/observability · high confidence
New observability metrics for ingestion validation, error tracking, and storage health
The backend now exposes several new OpenTelemetry metrics to improve visibility into ingestion behavior and system health. The \opik.ingestion.uuid\_v7.rejected\ counter tracks UUIDv7 validation outcomes (audit, reject, or bypass) to support the UUIDv7 re-enablement campaign. The \opik.errors.count\ counter provides error visibility broken down by root cause type, HTTP method, and endpoint, complementing standard HTTP status codes. Server-side ingestion size guards now emit \ingestion\_size\_guard\_rejections\_total\ to track requests rejected due to size limits. Additionally, new ClickHouse partition-health metrics report per-table and per-partition storage statistics (rows, bytes, parts) and lightweight-deleted row counts to aid in database maintenance.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/metrics · high confidence
New private API endpoints for Agent Configs, Insights, Alerts, and Annotation Queues
The backend now exposes a suite of new private REST endpoints under \/v1/private/\ to support advanced agent and workspace management capabilities. These include \AgentConfigsResource\ for managing optimizer configurations and blueprints, \AgentInsightsResource\ and \AgentInsightsJobsResource\ for generating and tracking agent performance reports, \AlertResource\ for creating and managing workspace-level alerts, and \AnnotationQueuesResource\ for handling human-in-the-loop annotation workflows. Additionally, \AssertionResultsResource\ provides batch ingestion for evaluation assertions, and \AttachmentResource\ handles file uploads and downloads for traces and spans. These endpoints enable programmatic control over agent optimization, observability, and data annotation features.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/priv · high confidence
New project and workspace metrics API request/response models
The backend now exposes structured request and response types for project and workspace-level metrics, enabling users to query metrics such as trace/span/thread counts, durations, token usage, costs, and error rates. This includes support for grouping (breakdown) by dimensions like tags, metadata, name, error info, model, provider, and guardrail name, with validation ensuring field compatibility with specific metric types. New KPI card endpoints provide current and previous period values for counts, errors (as percentages), average duration, and total cost, while workspace-level endpoints allow aggregating span metrics and token usage names across multiple projects.
apps/opik-backend/src/main/java/com/comet/opik/api/metrics · high confidence
New self-host configuration documentation for usage statistics, dataset migration, CSV uploads, LLM registry, and UUID validation
The self-host configuration documentation has been expanded with five new pages. Anonymous usage statistics reporting is now documented, including the data collected and the \OPIK\_USAGE\_REPORT\_ENABLED\ toggle. A comprehensive guide covers the automated dataset versioning migration introduced in Opik 1.9.92, detailing the Liquibase and MySQL counter steps and relevant environment variables. Large CSV upload support (up to 2GB) is documented with configuration for server timeouts, Nginx limits, and batch processing. The LLM model registry page explains how to configure the backend to fetch models from a remote CDN or local overrides. Finally, UUIDv7 ingestion validation is documented, describing the opt-in window check that rejects client-supplied IDs with out-of-range timestamps to protect partition layout.
apps/opik-documentation/documentation/fern/docs-v2/self-host/configure · high confidence
New self-hosted documentation site and migration guide
The self-hosted documentation has been restructured into a new v2 site under \documentation/fern/docs-v2/self-host\, introducing a comprehensive Platform Architecture overview, a dedicated Helm Chart Migration guide for moving from Bitnami to official images, and an Opik 2.0 migration guide for upgrading from version 1.x. The new site also includes updated guides for Kubernetes and local Docker Compose deployments, a self-host changelog, and detailed sections on ClickHouse backup, scaling, and troubleshooting.
apps/opik-documentation/documentation/fern/docs-v2/self-host · high confidence
New session redirect endpoints for projects, datasets, experiments, and optimizations
A new REST resource at /v1/session/redirect has been added to handle redirects for SDK-generated links. It exposes four GET endpoints—/projects, /datasets, /experiments, and /optimizations—that accept trace\_id, dataset\_id, experiment\_id, optimization\_id, workspace\_name, and a base64-encoded path query parameter, then return HTTP 303 redirects to the corresponding destination URLs.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/session · high confidence
New shared charting components for legends and tooltips
The frontend now includes a new set of reusable charting components under \src/shared/Charts\, including \ChartHorizontalLegend\, \ChartVerticalLegend\, \ChartTooltipContent\, \LegendItem\, \LineChart\, and \RadarChart\. These components provide consistent, interactive visualization features such as horizontal and vertical legends with line-highlighting on hover, customizable tooltips with header and value rendering, and support for area and radar chart types. This refactoring centralizes chart UI logic, improving consistency and maintainability across the application's data visualization features.
(repo-wide) · high confidence
New webhook event data models for alerts and metrics
The backend now introduces specific data structures to handle alert notifications sent via webhooks. A new \AlertEvent\ record serves as the initial evaluation payload, carrying workspace, user, and project context along with the alert type. For metrics-based alerts, \MetricsAlertPayload\ exposes threshold evaluation results, including support for multiple conditions within an AND-group via a \conditions\ list, while maintaining backward compatibility with scalar fields. These models are consumed by the generic \WebhookEvent\ class, which manages the delivery details such as the target URL, retry logic, and authentication secrets, ensuring that alert triggers are correctly formatted and routed to external endpoints.
apps/opik-backend/src/main/java/com/comet/opik/api/events/webhooks · high confidence
OpenRouter LLM provider integration
The backend now supports the OpenRouter provider for LLM interactions. This change introduces the OpenRouter-specific error handling (OpenRouterErrorMessage), registers the OpenRouter service provider (OpenRouterLlmServiceProvider) to route requests via the OpenAI-compatible client generator, and defines the comprehensive list of available models (OpenRouterModelName) sourced from OpenRouter's documentation.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/openrouter · high confidence
OpenTelemetry trace ingestion now supports JSON payloads
The backend infrastructure in the OpenTelemetry package now includes a new JSON message body reader alongside the existing Protobuf reader. This allows the system to ingest OpenTelemetry trace data sent as JSON (application/json) in addition to the previously supported Protobuf format, enabling clients that cannot easily serialize to Protobuf to still send trace data to the backend.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/otel · high confidence
Operator runbook and drivers for the spans table cutover to partitioned v2
This change introduces the operational tooling and documentation required to migrate the large, unpartitioned ClickHouse \spans\ table to the new \spans\_local\_v2\ schema (weekly-partitioned, denullified, and \is\_deleted\-ready). It includes a comprehensive runbook (\README.md\) detailing a near-zero-downtime cutover strategy involving backfill, delta replay, and atomic exchange, alongside executable shell drivers (\backfill.sh\, \delta\_replay.sh\, etc.) and reference SQL scripts (\000001\_backfill\_spans\_local\_v2.sql\, \000002\_delta\_and\_deletion\_replay.sql\). The drivers are specifically tuned for the scale of the spans table, addressing large active parts, wide uncompressed rows, and high partition spread by adjusting thresholds for memory usage, disk headroom, and partition limits, ensuring data integrity during the transition.
apps/opik-backend/data-migrations/spans-local-v2-cutover · high confidence
Opik backend container image and configuration scaffolding
The \apps/opik-backend\ module now ships a self-contained Docker image based on Amazon Corretto 25 (Java 25) and AL2023. The image includes a pre-configured AWS RDS certificate bundle for secure MySQL connections, an OpenTelemetry Java agent for metrics, and a dedicated read-only ClickHouse user for Agent Insights (gated by the \TOGGLE\_OLLIE\_ENABLED\ flag). It also introduces a \rebaseline\_db\_changelog.sh\ utility to repair lost Liquibase ledgers without re-executing migrations, and a \config.yml\ that exposes granular toggles for the traces/spans Distributed table wrap, UUIDv7 ingestion validation, and ClickHouse async-insert tuning.
apps/opik-backend · high confidence
Optional usage limit enforcement for API requests
The backend now supports optional enforcement of usage quotas (such as span counts) on API endpoints. By annotating methods with @UsageLimited, the system checks the user's current usage against configured limits; if a limit is exceeded, the request is rejected with a 402 Payment Required status and a configurable error message. This feature is controlled via the 'usageLimit' configuration, allowing administrators to enable or disable quota checks as needed.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/usagelimit · high confidence
REST API overview documentation added for Opik Cloud and Open-Source
The REST API reference now includes an overview page that distinguishes between the Open-Source platform and Opik Cloud. Users can find specific usage instructions for each, including the localhost URL for local development and the required authentication headers (API Key and Workspace) for Opik Cloud or on-premise Comet installations, with a note that the Bearer prefix is not used.
apps/opik-documentation/documentation/fern/docs-v2/reference/rest-api · high confidence
Standardized development environment with unified pre-commit hooks and AI agent configuration
The repository now includes a comprehensive set of configuration files to standardize the local development experience. A unified \.pre-commit-config.yaml\ enforces consistent code quality, linting, and security checks (including SQL string-formatting detection and GitHub Actions security scanning) across all modules. New \.editorconfig\ and \.java-version\ files ensure consistent formatting and specify Java 25. AI coding assistants are supported via \.cursorignore\, \AGENTS.md\, and symlinks to \.agents\. Additionally, \.env.template\ provides a reference for required environment variables, and \.git-blame-ignore-revs\ helps maintain clean history by ignoring large, non-functional reformatting commits.
(repo-wide) · high confidence
Support for OpenAI Responses API and improved quota error handling
The OpenAI LLM provider now supports the modern OpenAI Responses API alongside the traditional Chat Completions API, allowing users to select the API pipeline mode in their provider configuration. This change introduces new backend components to translate between the proxy's Chat Completions wire format and the Responses API, including handling for structured output strictness and streaming. Additionally, the provider now correctly treats 'insufficient\_quota' (HTTP 429) errors as non-retryable, preventing error storms when an API key is out of credits, and improves error mapping to surface specific diagnostic codes from OpenAI.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/openai · high confidence
Support for video, audio, and file content in custom LLM models
The backend now supports sending video, audio, and file attachments to custom LLM models via the LangChain4j integration. New model wrappers (OpikOpenAiChatModel, OpikGeminiChatModel) and message types (OpikUserMessage, OpikContent) handle multi-modal inputs, including automatic conversion of video content to image format for Gemini compatibility and direct URL support for audio and video in OpenAI-compatible APIs.
apps/opik-backend/src/main/java/com/comet/opik/domain/llm/langchain4j · high confidence
Traces cutover migration scripts for app analytics
Added the SQL runbook and driver scripts for the traces-local-v2 cutover in the app analytics database. This includes the backfill (000001), delta-insert and deletion-replay (000002), and the exchange-and-wrap (000003) steps that migrate trace data to a new partitioned schema and wrap it with a Distributed table for sharding readiness. It also provides the rollback procedures (000004), covering reverse-replay of deletes, sentinel repair for non-nullable columns, and stage-specific swaps (discard shadow, exchange back, promote original, unwrap) to restore the previous state or recover from wrap failures.
apps/opik-backend/data-migrations/traces-local-v2-cutover/scripts/db-app-analytics · high confidence
Traces cutover migration tooling and scripts
Introduces the operational scripts and SQL templates for the traces cutover migration, enabling the transition from the legacy \traces\ table to the new \traces\_local\_v2\ schema. The suite includes \backfill.sh\ for copying historical data in time-sliced, resumable batches; \delta\_replay.sh\ for replaying recent changes; \estimate.sh\ for projecting migration duration; and \exchange\_and\_wrap.sh\ for the atomic table swap and Distributed table topology update. These tools provide operators with configurable controls for throughput, resource usage, and safety gates to execute the data migration with minimal downtime.
apps/opik-backend/data-migrations/traces-local-v2-cutover/scripts · high confidence
Security
Upgrade frontend Docker base image to fix [CVE redacted]
The frontend Docker build image has been upgraded to patch the nghttp2-libs vulnerability ([CVE redacted]), improving the security posture of the deployed application.
apps/opik-frontend · high confidence
Behavioural changes
Backend sorting logic now supports dynamic keys and specific null handling for experiment scores
The backend introduces a new SortingQueryBuilder component that enables sorting on dynamic fields (such as JSON metadata or experiment scores) by generating SQL with bound parameters and handling null values explicitly. For experiment scores specifically, the logic now uses a pre-computed aggregation alias and checks for key existence to determine null direction, ensuring that missing scores are handled consistently during sort operations.
apps/opik-backend/src/main/java/com/comet/opik/domain/sorting · high confidence
Cursor extension introduces structured trace processing and deduplication
The Cursor extension now reads conversation data directly from the local SQLite database and processes it into structured Opik traces. This change adds logic to classify chat bubbles (user, tool, error, thinking) and order them correctly, even when branches are edited away. It implements a request ledger to deduplicate traces and prevent duplicate uploads across edits, using deterministic UUIDv7 identifiers. The extension also handles usage enrichment, ensuring token counts and costs are accurately attributed to the correct traces and revisions.
extensions/cursor · high confidence
Data migration for dataset item \`data\` field structure
A new data migration script (1.0.3) is provided to backfill the \data\ field in existing dataset item rows to match the new dynamic fields structure. This migration updates legacy rows by merging \input\, \expected\_output\, and \metadata\ into the \data\ map column, ensuring consistency for installations that generated datasets prior to this release. Users must manually execute the provided ClickHouse SQL scripts outside peak hours to avoid resource contention.
apps/opik-backend/data-migrations/1.0.3 · high confidence
Enable Redis ElastiCache IAM authentication and fix MinIO compatibility
The backend now supports authenticating to AWS ElastiCache for Redis using IAM credentials via a new credential resolver, allowing users to connect to secured Redis clusters without managing static passwords. Additionally, the S3 client configuration for MinIO has been updated to disable SHA-256 checksum validation, resolving compatibility issues with MinIO deployments that do not support this validation method.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/aws · high confidence
Fix StringTemplate memory leak and add template error logging
The backend now uses a new TemplateUtils factory to create StringTemplate instances, preventing memory leaks caused by templates being cached in a shared singleton group. This change ensures each template is ephemeral and garbage-collectible after use. Additionally, a new TemplateLogger has been introduced to capture and log template compilation and runtime errors, while suppressing specific debug-level warnings for undefined attributes in conditional SQL fragments.
apps/opik-backend/src/main/java/com/comet/opik/utils/template · high confidence
Improved LLM error handling and multimodal support in chat completions
The backend now introduces a centralized ChatCompletionService that normalizes requests and handles errors more robustly: it propagates the provider's actual HTTP status code to the client instead of masking failures as generic 500 errors, and it stops retrying calls that fail due to unsupported features. Additionally, the system now supports multimodal inputs by expanding image placeholders in user messages into structured content for vision-capable models, and it enforces Anthropic-specific sampling constraints (such as excluding temperature and top\_p together) at the routing layer.
apps/opik-backend/src/main/java/com/comet/opik/domain/llm · high confidence
Improved OpenTelemetry provider mapping and cost attribution
The backend now correctly maps OpenTelemetry GenAI semantic-convention provider names to Opik's canonical provider identifiers, ensuring that spans from integrations like Vertex AI, AWS Bedrock, and Azure are attributed to the correct pricing tables. This change resolves issues where unmapped or ambiguous provider values resulted in zero-cost spans or incorrect grouping. Additionally, new mapping rules and resolvers have been added to support specific integrations including Claude Code, Elastic Inference Service, LiveKit, and PydanticAI, improving the accuracy of input, output, and usage data captured from these sources.
apps/opik-backend/src/main/java/com/comet/opik/domain/mapping/otel · high confidence
Improved parameter parsing and JSON upload format handling
The backend now includes new infrastructure components to handle date/time parameters and JSON upload formats more robustly. Instant and LocalDate parameters are now parsed with specific error handling, supporting ISO-8601 formats and milliseconds since epoch for Instants, while providing clear error messages for invalid inputs. Additionally, a new MessageBodyReader ensures that JsonUploadFormat values from multipart form data are correctly parsed regardless of case sensitivity, addressing potential issues with Jersey's default enum provider.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/web · high confidence
Introduce pluggable template parser architecture with unescaped Mustache rendering
The backend now uses a new \TemplateParser\ interface with concrete implementations for Mustache, Python-style format strings, and Jinja2. This change introduces unescaped rendering for Mustache templates to ensure that JSON structures and special characters (like \=\ and backticks) are passed to LLM judges exactly as written, matching the frontend preview and Python parser behavior. It also adds robust error handling for malformed templates during variable extraction and provides a stub for future Jinja2 support.
apps/opik-backend/src/main/java/com/comet/opik/domain/template · high confidence
Introduce reactive webhook delivery with Redis stream decoupling
The backend now decouples webhook triggering from alert generation by publishing events to a Redis stream via the new WebhookPublisher, which is then consumed and sent asynchronously by the new WebhookHttpClient. This change enables reliable, retry-capable delivery of webhook notifications (including those for Slack and PagerDuty integrations) with configurable timeouts, destination validation, and encrypted secret handling, ensuring that alerting logic is not blocked by external HTTP response times.
apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/webhooks · high confidence
Introduces Opik Free Model provider with reasoning model constraints
The backend now includes a dedicated 'Opik Free' LLM provider that maps the generic 'opik-free-model' identifier to a specific underlying model configured in the system. This provider enforces specific behavioral constraints for reasoning models: it clamps the temperature to a minimum of 1.0 if a lower value is requested, and it hardcodes the reasoning effort to 'minimal' to reduce latency and cost for simple evaluation tasks. The provider is registered via a new Guice module and handles both standard and streaming chat completions while treating insufficient quota errors as non-retryable.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/freemodel · high confidence
Introduces a unified sorting framework for backend list endpoints
The backend now uses a centralized sorting infrastructure (SortingFactory, SortableFields, and SortingField) to standardize how list endpoints handle sorting. This change introduces specific sorting factories for key resources, enabling users to sort lists of Traces, Spans, Experiments, Datasets, Prompts, Prompt Versions, Projects, Dashboards, Alerts, Annotation Queues, Automation Rules, and Agent Insights Issues by their respective columns (such as name, timestamps, duration, cost, and status). The new system also supports sorting by dynamic JSON fields (e.g., feedback scores) and ensures that invalid or unsupported sort requests are handled gracefully without causing server errors.
apps/opik-backend/src/main/java/com/comet/opik/api/sorting · high confidence
Introduces structured request/response models and robust error handling for Python evaluators
The backend now uses dedicated record classes (PythonEvaluatorRequest, PythonEvaluatorResponse, PythonEvaluatorErrorResponse, PythonScoreResult) to serialize and deserialize Python evaluation payloads, ensuring snake\_case JSON mapping and ignoring unknown fields. PythonEvaluatorService implements non-blocking, buffered response processing with retry logic and null-safe error extraction, preventing NullPointerExceptions when the Python evaluation service returns empty or malformed bodies and surfacing truncated error messages to the user.
apps/opik-backend/src/main/java/com/comet/opik/domain/evaluators/python · high confidence
New API request and response models for Connect session management
The backend now exposes specific data structures for the Connect API, introducing ActivateRequest, CreateSessionRequest, and CreateSessionResponse. These models define the expected input for activating runners (requiring a runner name and HMAC tag) and creating sessions (requiring a project ID, activation key, optional TTL, and runner type), as well as the resulting session and runner identifiers returned to the client.
apps/opik-backend/src/main/java/com/comet/opik/api/connect · high confidence
New API request validation rules for dataset, experiment, and provider operations
The backend now enforces stricter validation on several API request types to prevent invalid configurations and improve error reporting. Dataset item batch updates and deletions now require either specific item IDs or filter criteria (with a required dataset ID for filters), but never both simultaneously. Experiment bulk uploads validate that item-level trace projects match the request-level project name and reject records containing both evaluate task results and traces. Provider configuration validation ensures API keys are present when required, prevents mixing API keys with auth configs, and enforces naming requirements for custom LLM and Bedrock providers. Additional validation covers metric breakdown configurations (requiring sub-metrics for duration, feedback, and token usage metrics), duration precision (seconds or higher), and execution policy thresholds.
apps/opik-backend/src/main/java/com/comet/opik/api/validation · high confidence
New Semgrep rules to prevent SQL injection and datetime wrapping bugs
Added two new Semgrep pre-commit hooks for Java SQL code: \sql-query-clause-splice\ (ERROR) blocks SQL strings using \%s\ for clause splicing to prevent injection, and \sql-narrow-datetime-on-id\ (ERROR) flags narrow datetime conversions on UUIDv7-derived IDs to prevent far-future date wrapping. A secondary reporting rule, \sql-query-format-slot\ (WARNING), alerts on any other \%s\ format slots in SQL literals. These rules apply to production code in \apps/opik-backend/src/main/java\ and exclude test sources.
.semgrep · high confidence
New agent rules for deployment, code style, and security
The repository now includes structured agent rules in \.agents/rules/\ that guide development workflows. The \agents.mdc\ file defines domain routing (e.g., backend vs. frontend skills) and mandates the use of the V2 frontend structure, noting the removal of V1. The \code-style.mdc\ file enforces concise commenting and prohibits over-engineering. The \git-workflow.mdc\ file standardizes branch naming, commit formats, and PR body structures, including specific conventions for Jira key references to prevent unintended linking. Finally, the \security.mdc\ file explicitly forbids building SQL from formatted strings and requires parameterized queries, alongside general rules against hardcoding secrets.
.agents/rules · high confidence
New error handling infrastructure with structured metrics and validation
The backend introduces a new set of exception classes and mappers in the error package to standardize API error responses and improve observability. This includes specific exceptions for conflicts (ConflictException, EntityAlreadyExistsException, IdentifierMismatchException) and a new InvalidUUIDException for UUIDv7 ingestion validation, which maps to HTTP 400 errors. A JsonProcessingExceptionMapper now distinguishes between structural JSON errors (400) and size-limit violations (413), recording metrics for ingestion size guards. Additionally, generic exception mappers (OpikGenericExceptionMapper, OpikWebApplicationExceptionMapper) have been added to count 5xx server errors under the opik.errors.count metric, while maintaining existing logging behavior.
apps/opik-backend/src/main/java/com/comet/opik/api/error · high confidence
Online evaluation runs are persisted as hidden monitoring traces
Online evaluation runs (LLM-as-judge scoring) are now recorded as hidden monitoring traces and spans in the backend. This ensures that evaluation activity is tracked for observability and cost accounting without cluttering the default traces view or triggering recursive self-evaluation.
apps/opik-backend/src/main/java/com/comet/opik/domain/evaluation · high confidence
Opik documentation v2 site launch with new core pages
The documentation site has been updated to version 2, introducing a new home page, a quickstart guide for Python and TypeScript SDKs, and a dedicated FAQ page covering rate limits, size limits, and API keys. A new 'Ollie' page details the AI assistant and its integration via \opik connect\, while an upgrade guide explains the shift to project-scoped data in Opik 2.0. The reference overview now explicitly lists the Ruby OpenTelemetry SDK alongside Python, TypeScript, and REST API docs.
apps/opik-documentation/documentation/fern/docs-v2 · high confidence
Python sandbox executor optimized for faster scoring with bytecode compilation and import patching
The Python sandbox executor now significantly reduces cold-start latency for scoring tasks. The Docker image is built using a multi-stage process that compiles standard library and dependency bytecode at build time, ensuring it is available at runtime without recompilation. Additionally, the scoring runner implements a lightweight import-patching mechanism that delays the loading of the heavy 'opik' package until it is actually needed, preventing timeout issues caused by slow initial imports under CPU contention.
apps/opik-sandbox-executor-python · high confidence
Refactored Anthropic provider integration to support new Claude models and adaptive thinking
The Anthropic LLM provider implementation has been rewritten to support the latest Claude model family (including Sonnet 4.5, Opus 4.5/4.6/4.7/4.8, and Sonnet 5) and to correctly handle Anthropic's adaptive thinking feature. The new client generator and mapper logic automatically gates temperature and top\_p sampling parameters for models that reject them (such as Sonnet 5 and Opus 4-7/4-8) and ensures that \max\_tokens\ is always explicitly set and sufficiently large to accommodate extended thinking budgets, preventing API 400 errors. This change also standardizes the provider service registration and improves error mapping for Anthropic-specific exceptions.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/antropic · high confidence
Refactored Gemini provider infrastructure with structured thinking and video support
The Gemini LLM provider implementation has been restructured into a modular set of components within the backend infrastructure. This change introduces explicit support for the Gemini 'thinking' configuration, allowing users to control reasoning depth via thinking levels or budgets, while ensuring that internal reasoning traces are suppressed from the final response text to match expected model behavior. Additionally, the provider now handles video content by automatically converting video URLs into image content compatible with the Gemini API, including MIME type detection. The update also expands the supported model registry to include recent Gemini versions (such as 2.5, 3, and 3.1 series) and refactors error handling and client generation for better maintainability.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/gemini · high confidence
Refactored authentication infrastructure with Redis caching and granular permissions
The authentication layer in the backend has been restructured to support more flexible and performant access control. A new Redis-based caching service (AuthCredentialsCacheService) now stores and resolves API key credentials, workspace metadata, and user permissions, significantly reducing latency for repeated authentication checks. The authentication flow now distinguishes between local installations (using a default workspace) and remote services, while also introducing support for CIPX device tokens restricted to specific ingest endpoints. Additionally, a new annotation-driven permission system (RequiredPermissions) allows endpoints to declare specific workspace permissions, which are resolved and enforced via the new AuthDynamicFeature and AuthFilter components.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/auth · high confidence
Refactored backend filtering infrastructure to support dynamic fields and robust JSON path handling
The filtering logic in the analytics backend has been restructured to introduce a strategy-based approach for building queries across all entity types (traces, spans, experiments, datasets, prompts, etc.). This change adds comprehensive support for filtering on dynamic fields (such as metadata and custom dataset columns) by correctly resolving JSON paths for ClickHouse, ensuring that special characters in keys are handled safely. It also implements sentinel-aware filtering for time-based fields (end\_time, duration, ttft) to correctly exclude rows with missing or null values, and introduces a new ENUM\_LEGACY field type to maintain compatibility with legacy source data while standardizing filter behavior.
apps/opik-backend/src/main/java/com/comet/opik/domain/filter · high confidence
Refactored project stats mapping and merging logic
The backend now uses a dedicated StatsMapper to map database rows into ProjectStats objects, explicitly handling fields such as thread count, guardrails failed count, and error counts (recent, past period, and total). A new StatsMerger component combines trace/span statistics with separate feedback-score statistics, ensuring that feedback data is spliced into the correct position within the stats list without resurrecting projects that lack trace data. This change centralizes the mapping logic and clarifies how different statistical aggregates are combined for project-level views.
apps/opik-backend/src/main/java/com/comet/opik/domain/stats · high confidence
SDK configuration flow improvements
The SDK's configuration process has been enhanced to provide a smoother onboarding experience. The \opik configure\ command now intelligently suggests the most recently created project during setup, and the configuration flow has been hardened to better handle interactive environments and API key validation, ensuring that users are guided more effectively when setting up their workspace and API credentials.
sdks · high confidence
Structured output now uses native JSON response formats where supported
The backend now automatically selects the most reliable method for obtaining structured JSON from LLMs based on provider capabilities. For providers that support native JSON response formats (OpenAI, Gemini, OpenRouter, and Vertex AI), the system now uses the provider's built-in JSON schema enforcement via the \ToolCallingStrategy\. For providers that do not support this feature (Anthropic, Bedrock, Ollama, and custom LLMs), it falls back to appending detailed text instructions to the user message via the new \InstructionStrategy\. This change improves reliability and reduces parsing errors for supported models while maintaining compatibility with others.
apps/opik-backend/src/main/java/com/comet/opik/domain/llm/structuredoutput · high confidence
Vertex AI integration now properly manages client lifecycle and supports new Gemini models
The Vertex AI backend infrastructure has been refactored to prevent thread leaks by ensuring the Google GenAI client is created and closed per request, addressing previous stability issues. This change also introduces support for a broader range of Gemini models (including Gemini 2.5, 3, and 3.5 variants) and enables native thinking-level configuration for Vertex Gemini 3. Additionally, a workaround is implemented to handle requests that lack user or AI messages, preventing 500 errors during generation.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/llm/vertexai · high confidence
Workspace metadata tracking and legacy feedback scores support
The backend now tracks workspace-level metadata, including the timestamp of the first reported trace and a flag indicating whether legacy feedback scores exist. This enables the system to correctly handle legacy data during stats queries by conditionally including the legacy feedback\_scores table, and ensures accurate audit trails for workspace initialization.
apps/opik-backend/src/main/java/com/comet/opik/domain/workspaces · high confidence
Fixes
Exclude demo project usage from billing and BI reports
Usage reports and daily billing counts now correctly exclude demo projects. Previously, demo projects (created per signup) caused query performance issues and inaccurate counts because their IDs were unbounded. This change moves the exclusion logic from SQL queries to Java, filtering out demo project IDs after aggregation to ensure accurate workspace and user-level usage totals for billing and business intelligence.
apps/opik-backend/src/main/java/com/comet/opik/domain/utils · high confidence
Fix HTTP 500 errors for error message serialization in streaming endpoints
Added custom JAX-RS MessageBodyWriters for \ErrorMessage\, \InternalErrorMessage\, \ValidationErrorMessage\, and \JsonNode\ types to ensure they are correctly serialized to JSON when the response content type is \application/octet-stream\. This prevents HTTP 500 errors that previously occurred in streaming endpoints when the framework could not find a suitable writer for these error types.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/json · high confidence
Fixes outbound HTTP double-decompression and adds connection pool health monitoring
The backend now prevents a ZipException caused by Apache HttpClient and Jersey both decompressing gzip responses by stripping stale Content-Encoding headers after HttpClient's transparent decode. Additionally, a new liveness health check detects when the shared outbound HTTP connection pool has shut down, ensuring Kubernetes restarts the pod to restore connectivity rather than leaving it serving failed requests.
apps/opik-backend/src/main/java/com/comet/opik/infrastructure/http · high confidence
Test coverage
Add trace test assertion utilities; Added MySQL query plan-shape regression tests; Added Podam test data factories for backend unit tests; Added Podam test data manufacturers for backend API models; Added end-to-end test for MySQL RDS IAM authentication; Added end-to-end tests for CORS configuration; Added integration and unit tests for MCP OAuth client lifecycle and security; Added integration test for project metadata cache behavior; Added integration tests for the backend analytics and usage reporting infrastructure; Added span test assertion utilities; Added test coverage for backend utility classes; Added test fixtures for topology-aware traces DDL pattern validation; Added test infrastructure and validation for test timeouts; Added test utility clients for backend API resources; Added tests for Anthropic client generation and parameter mapping; Added tests for AttachmentUploadRequested event; Added tests for BreakdownQueryBuilder input validation; Added tests for ClickHouse partition metrics, error classification, and ingestion size guards; Added tests for DestinationGuard validation logic; Added tests for FreeModel LLM provider temperature clamping and parameter forwarding; Added tests for Gemini LLM infrastructure mappers and video support; Added tests for HTTP client gzip handling and health checks; Added tests for LLM infrastructure components; Added tests for LLM service error handling, content normalization, and sampling parameters; Added tests for LLM-as-Judge evaluator filter deserialization and multimodal message serialization; Added tests for MCP OAuth resource endpoints and bundle wiring; Added tests for OpenAI LLM provider infrastructure; Added tests for OpikUserMessage serialization; Added tests for Optimization Log Sync Service; Added tests for Redis infrastructure reliability and configuration; Added tests for SortingFactory null/blank field handling and server-side bind key generation; Added tests for TemplateUtils memory leak prevention; Added tests for UUIDv7 validation, zero-rows retry, and ClickHouse health checks; Added tests for Vertex AI client lifecycle and configuration; Added tests for alert debouncing and Redis stream configuration; Added tests for annotation queue locking, dataset export, and assertion mapping; Added tests for authentication infrastructure components; Added tests for cost calculation logic; Added tests for custom LLM provider infrastructure; Added tests for evaluator domain logic and online scoring infrastructure; Added tests for event processing components; Added tests for experiment aggregation and denormalization logic; Added tests for filter query builder sentinel handling, v2 client integration, and JSON path utilities; Added tests for internal analytics and usage endpoints; Added tests for rate limiting infrastructure; Added tests for retention policy enforcement and velocity estimation; Added tests for the Is Alive health check resource; Added tests for the cache infrastructure; Added tests for the session redirect resource; Added tests for traces DDL topology guards and changelog re-baseline; Added unit and integration tests for backend job components; Added unit tests for LLM usage extraction and online evaluation recording; Added unit tests for Mustache and Python template parsers; Added unit tests for Ollie state service upload, download, and delete operations; Added unit tests for OpenTelemetry mapping utilities; Added unit tests for Python Evaluator Service; Added unit tests for StatsMapper project and score statistics; Added unit tests for ValidationErrorMessageBodyWriter; Added unit tests for attachment stripping and reference extraction; Added unit tests for backend API configuration and data models; Added unit tests for backend error mappers; Added unit tests for the pytest plugin hooks; Added validation script for dataset versioning migration; Added validation tests for backend API constraints; Centralized dataset item test assertions; Expanded test coverage for backend infrastructure components; Expanded test infrastructure and utilities for backend resource tests; Test infrastructure updates for ClickHouse and JUnit timeouts.
Dependencies
Opik Backend dependency baseline established
The Opik backend build configuration (pom.xml) is introduced, defining the initial dependency versions for the service. This includes upgrading the core framework to Dropwizard 5.0.2, setting the OpenTelemetry version to 2.31.1, and pinning key libraries such as Redisson 4.7.0, ClickHouse JDBC 0.10.0, and the AWS SDK BOM 2.54.2.
(dependencies) · high confidence
Housekeeping
Weekly changelog updates (September 2026)
The documentation site now includes weekly changelog entries for the weeks of August 10, August 24, August 31, and September 7, 2026, keeping the public release history current.
apps/opik-documentation/documentation/fern/docs-v2/changelog · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 53.
Lenses
- Code Health 81
- Architecture 64
- Maturity 77
- Readiness 52
- Security 57
- Accessibility 47
Changes since last survey
- 300 commits — 220 feature/other, 80 fixes
By area
- apps/opik-backend — 124 commits
- sdks/python — 36 commits
- sdks/typescript — 36 commits
- tests_end_to_end/e2e — 35 commits
- apps/opik-frontend — 23 commits
- apps/opik-documentation — 14 commits
- .github/workflows — 13 commits
- (root) — 5 commits
- tests_end_to_end/coverage — 5 commits
- extensions/cursor — 4 commits
- .agents/commands — 1 commit
- .agents/skills — 1 commit
- apps/opik-python-backend — 1 commit
- deployment/helm_chart — 1 commit
- scripts/dev-runner-platform.sh — 1 commit
Notable commits
- fix: [issue-8175] [SDK] fix: delegate ConversationThreadMetric.ascore to score instead of raising (#8176)
- fix: [NA] Fix the prod nightly's chronic e2e failures: id-window seeds, SPA shell wait, reason poll (#8301)
- fix: [NA] [BE] fix: bind dataset item filters to item-level columns in page and count (#8040)
- fix: [NA] [BE] fix: register cerebras as a canonical provider so Cerebras model prices load (#7731)
- fix: [NA] [BE] fix: register deepinfra as a canonical provider so DeepInfra model prices load (#7715)
- fix: [NA] [BE] fix: register snowflake as a canonical provider so Snowflake Cortex model prices load (#7714)
- fix: [NA] [BE] fix: stop wrapped week bounds dropping far-future rows from trace reads (#8096)
- fix: [NA] [BE] fix: validate webhook destinations and trim test error detail (#8354)
- fix: [NA] [DEV] fix: point EM backend S3 at the published MinIO API port (#8111)
- fix: [NA] [EXT] fix: include cache tokens in reported prompt and total tokens (#8011)
- fix: [NA] [EXT] fix: prevent duplicate Cursor traces across edits (#8090)
- fix: [NA] [FE] fix: round the seconds remainder in formatDuration (#8248)
- fix: [NA] [FE] fix: stop the model-parameters panel offering what it cannot deliver (#8223)
- fix: [NA] [QA] fix: gate Playground dataset runs on the items query resolving (#8300)
- fix: [NA] [QA] fix: make the e2e tsc gate actually type-check the estate (#8279)
- fix: [NA] [QA] fix: stop waiting on the suite verdict in waitForRunsComplete (#8304)
- fix: [NA] [SDK] docs: document entrypoint/environment args and fix garbled task_threads docstring (#8165)
- fix: [NA] [SDK] fix: keep anthropic and openai tracking working after upstream releases (#8309)
- fix: [NA] [SDK] fix: offer to remove the uv tool install that hijacks uvx opik-mcp (#8313)
- fix: [NA] [SDK] fix: reject duplicate items in SpearmanRanking (#7290)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
comet-ml/opik was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0647a9c7d04db693855cbda09cfe8e2f9daca786 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.