Skip to content
CAI
Software that uses CAICheck a score

gocolly/colly

64.3

Adequate · 6 August 2026

5k

lines of production code

Go

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Colly is a high-performance web scraping framework for Go that provides a comprehensive toolkit for building robust crawlers. It supports advanced features such as parallel and asynchronous scraping, request queuing, and rate limiting to manage concurrency and avoid overwhelming target servers. The system includes built-in extensions for user-agent rotation, URL filtering, and proxy switching, alongside a CLI for rapid project scaffolding. Additionally, it offers debugging tools and in-memory storage to facilitate development and maintenance.

How it got here

2017 — API expansion and example library growth

19 changes.

This period focused on expanding the Colly framework with new features such as HTTP tracing, proxy rotation, and a CLI code generator, alongside a significant addition of diverse web scraping examples. The release of version 2.1.0 introduced functional options for configuration and enhanced debugging capabilities, while removing outdated examples to streamline the codebase.

2018–2019 — Concurrency and extension features

8 changes.

This period focused on enhancing Colly's capabilities for concurrent crawling and request management by introducing a request queue and in-memory storage backend. The team also expanded the library's utility with an extensions package for common scraping tasks and provided new examples demonstrating these features.

Features

Add CLI tool to generate scraper templates

A new command-line interface has been added to the Colly framework, allowing users to generate starter Go code for scrapers. The \colly\ binary now supports a \new\ subcommand that scaffolds a new scraper project, with options to pre-configure allowed hosts and include common callback templates (HTML, request, response, and error handlers) directly in the generated code.

cmd · high confidence

Add Instagram scraper example

A new example script (\_examples/instagram/instagram.go) has been added that scrapes and downloads images from an Instagram profile page, handling pagination and API requests.

_\examples/instagram · high confidence

Add Open edX courses scraping example

Added a new Go example that demonstrates how to scrape course data from the IndonesiaX platform using the gocolly/colly/v2 library. The example defines a Course struct and uses functional options to configure the collector, including allowed domains and a local cache directory, then parses HTML to extract course details and outputs the results as JSON.

_\_examples/openedx\courses · high confidence

Add Reddit scraping example with async parallelism

Added a new example at \_examples/reddit/reddit.go that demonstrates how to use the colly library to scrape Reddit. The example shows how to configure the collector for async parallelism, set rate limiting with a random delay, and extract story titles and URLs from the old.reddit.com domain.

_\examples/reddit · high confidence

Add XKCD store scraping example

A new example has been added to demonstrate scraping the XKCD online store. The example uses the functional options pattern to configure the Colly collector, specifically allowing requests only to 'store.xkcd.com'. It extracts product details such as name, price, URL, and image URL, and writes the results to a CSV file.

_\_examples/cryptocoinmarketcap, \_examples/shopify\_sitemap, \_examples/xkcd\store · high confidence

Add error handling and request context examples

Added two new example files, \_examples/error\_handling/error\_handling.go and \_examples/request\_context/request\_context.go, demonstrating how to handle errors and manage request context in Colly.

_\_examples/error\_handling, \_examples/request\context · high confidence

Add multipart form upload example

A new example demonstrating how to upload multipart form data (including files) using the Colly library is now available. The example sets up a local HTTP server and uses Colly's PostMultipart functionality to send a form containing text fields and an image file.

_\examples/multipart · high confidence

Add parallel scraping example with controlled concurrency

A new example at \_examples/parallel/parallel.go demonstrates how to perform parallel web scraping using Colly's async mode. The example configures the collector to visit links asynchronously and explicitly limits the maximum parallelism to 2 concurrent requests using colly.LimitRule, ensuring the scraper behaves gently on the target site.

_\examples/parallel · high confidence

Add proxy switcher example demonstrating IP rotation

A new example at \_examples/proxy\_switcher demonstrates how to configure Colly to rotate between multiple SOCKS5 proxies using the RoundRobinProxySwitcher. The example shows how to set a proxy function on the collector and access the active proxy URL in the response handler, allowing users to see how to implement proxy rotation for requests.

_\_examples/proxy\switcher · high confidence

Add request queue example using Colly

A new example in \_examples/queue demonstrates how to use the Colly request queue with two consumer threads, showing how to add URLs and requests to a queue and run them concurrently.

_\examples/queue · high confidence

Add round-robin proxy switching for HTTP, HTTPS, and SOCKS5

The proxy package now includes a new RoundRobinProxySwitcher function that allows users to rotate through multiple proxy URLs (supporting http, https, and socks5 schemes) on each request. This enables load balancing or redundancy by cycling through a list of proxy endpoints.

proxy · high confidence

Add scraper server example

A new example demonstrating a Go HTTP server that scrapes a given URL and returns the status code and link counts as JSON. The server listens on port 7171 and uses the Colly library to parse HTML and extract link information.

_\_examples/hackernews\_comments, \_examples/scraper\server · high confidence

Add web and log debuggers for debugging collector events

A new \debug\ package is introduced, providing a \Debugger\ interface and two implementations: a \LogDebugger\ that writes events to STDERR, and a \WebDebugger\ that exposes a web UI at 127.0.0.1:7676 to monitor active and finished requests. The \Event\ struct is defined with \RequestID\ and \CollectorID\ fields, allowing users to track and debug collector activity through either logging or a web interface.

debug · high confidence

Added URL filtering example using functional options

Added a new example demonstrating how to use functional options with the colly.NewCollector to apply URL filters via regexp patterns, restricting which URLs are visited during scraping.

_\_examples/url\filter · high confidence

Added factba.se transcript scraper example

A new example script (\_examples/factba.se/factbase.go) has been added that uses the Colly library to scrape and download transcripts from factba.se, saving the data as JSON files.

_\examples/factba.se · high confidence

Added local files example

Added a new example demonstrating how to use Colly to parse local HTML files. The example includes a Go script (local\_files.go) that uses a custom HTTP transport to read from the file system, along with sample HTML files (index.html, child\_page/one.html, etc.) to illustrate the feature.

_\_examples/local\files · high confidence

Added login example demonstrating authenticated scraping

A new example in \_examples/login demonstrates how to perform an HTTP POST to authenticate and then scrape a protected page using the Colly library.

_\examples/login · high confidence

Added rate-limiting example with async and debugger support

A new example at \_examples/rate\_limit demonstrates how to configure Colly with asynchronous requests, attach a log-based debugger, and enforce a concurrency limit of two threads for domains matching '\httpbin.\'.

_\_examples/rate\limit · high confidence

Colly v2.1.0 release with HTTP tracing, new callbacks, and proxy fixes

The Colly framework releases version 2.1.0, introducing HTTP tracing support to measure connection and first-byte durations, and a new OnResponseHeader callback to inspect headers before the body is processed. The release also adds the Collector.CheckHead option, improves proxy handling, and fixes POST revisit checking. Additionally, the CHANGELOG.md, .codecov.yml, and other repository metadata files are added to the codebase.

(repo-wide) · high confidence

Introduce in-memory storage backend for collectors

A new \storage\ package is introduced, providing an in-memory implementation of the \Storage\ interface. This allows collectors to track visited URLs and manage cookies in memory without persisting data to disk. The \InMemoryStorage\ struct implements the \Storage\ interface, which includes methods for initializing storage, tracking visited requests, and managing cookies for a given host.

storage · high confidence

Introduce request queue for concurrent crawling

The library now includes a new \queue\ package that provides a \Queue\ type for managing a pool of concurrent HTTP requests. This feature allows users to control the number of active crawling threads and manage a queue of URLs or \colly.Request\ objects. The implementation includes an in-memory storage backend with a configurable maximum size, and integrates with the \colly\ collector to process requests in parallel. A test suite is also added to verify the queue's behavior under concurrent load.

queue · high confidence

New example demonstrating random delay with limit rules

Added a new example in the \_examples/random\_delay directory that shows how to use the LimitRule's RandomDelay feature with an async collector. This example illustrates configuring a random delay of 5 seconds for requests matching the \httpbin.\ domain glob, alongside setting a parallelism of 2 and attaching a debugger for visibility.

_\_examples/random\delay · high confidence

New extensions package with user-agent and URL filtering helpers

The \extensions\ package has been added to Colly, providing several helper addons for common scraping tasks. Users can now easily randomize the User-Agent header for desktop browsers (Firefox, Chrome, Edge, Opera) and mobile devices (Pixel, Nexus) on every request. Additionally, a new \URLLengthFilter\ extension is available to automatically abort requests that exceed a specified URL length limit, and a \Referer\ extension is included to automatically set the Referer header based on the previous response URL.

extensions · high confidence

New web scraping examples for Coursera and Google Groups

Added two new example scripts for web scraping: one for extracting course data from Coursera and another for scraping Google Groups threads. The Coursera example demonstrates using colly's cache directory and expiration features, while the Google Groups example shows multi-collector scraping with functional options. Both examples output their results as formatted JSON.

_\_examples/coursera\courses · high confidence

Behavioural changes

Removal of Wikipedia list scraping example

The example demonstrating how to scrape Wikipedia lists using the Colly library has been removed from the codebase. Users relying on this specific example for reference or implementation will no longer find it in the examples directory.

examples · high confidence

Update Colly examples to use functional options for collector configuration

The basic and max\_depth examples have been updated to use the functional options pattern for configuring the Colly collector, such as colly.AllowedDomains and colly.MaxDepth, providing a more flexible and readable way to set up scraping behavior.

_\_examples/basic, \_examples/max\depth · medium confidence

Dependencies

Updated Go dependencies and enabled Go modules

The project now uses Go modules, with go.mod and go.sum files added to manage dependencies. Key updates include upgrading golang.org/x/net to v0.47.0, google.golang.org/protobuf to v1.36.10, and various other libraries such as github.com/antchfx/xmlquery to v1.5.0 and github.com/nlnwa/whatwg-url to v0.6.2. These changes ensure the project uses the latest stable versions of its dependencies.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 64 → 64 (+0.0)
  • Rubric changed (rubric-2026.08.18 → rubric-2026.08.19) — scores are not directly comparable.

Lenses

  • Architecture 100 → 100 (+0.0)
  • Maturity 53 → 53 (+0.0)
  • Readiness 67 → 67 (+0.0)
  • Security 76 → 76 (+0.0)

Resolved (1)

  • Off-boarding risk: anonymized user #1

New (1)

  • Off-boarding risk: anonymized user #1

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

gocolly/colly was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 6 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 20d31482af5f754832a753f88517f3dfa61d921f — the exact code this score is about.
  • Scored under rubric-2026.08.19 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer latest.