Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
204 changes: 79 additions & 125 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,176 +1,130 @@
# Rainlytics

Self-hosted web analytics for AWS sites, built on CloudFront logs.
Rainlytics is self-hosted web analytics for sites served by Amazon CloudFront.

[rainlytics.com](https://rainlytics.com "Rainlytics documentation")
It collects most measurements from CloudFront access logs. Your pages need no analytics
JavaScript, no third-party script tag and no connection to an analytics provider. Raw logs and
derived reports stay in your AWS account.

Rainlytics runs the whole analytics pipeline inside your own AWS account. Most
of what it reports is derived from the CloudFront access logs your distribution
already writes. A measured page downloads no analytics JavaScript, opens no
extra connection, and resolves no extra hostname.
An optional browser module records SPA route changes, custom events, Core Web Vitals and JavaScript
errors. The module goes into your existing bundle and sends events to your own domain.

An optional beacon covers what an access log cannot see accurately, such as
route changes in a single-page app, Core Web Vitals, JavaScript errors and
events the site raises itself. It is bundled into the site's own JavaScript and
reports back through the site's own domain (no second host, no separate script
tag).
[Read the documentation](https://rainlytics.com) or follow [Getting started](docs/getting-started/)
to deploy the first working pipeline.

Everything runs on usage-priced AWS services, batched and precomputed on a
schedule rather than processed per request. Nothing in the pipeline is always
on, and a low-traffic site should cost cents a month.
## How it works

## What works today
CloudFront standard logging v2 writes access logs to S3. Rainlytics creates a Glue table over those
objects and uses partition projection, so there is no Glue crawler. Athena computes common
questions on a schedule. The results are stored as small JSON summaries in S3.

CDK constructs for the collection half of the pipeline. A distribution's access
logs land in an S3 bucket, partitioned and carrying the field set the rollups
will read, and a Glue table describes them for Athena.

```typescript
import {
CloudFrontLogDelivery,
LogBucket,
LogTable,
QueryWorkgroup,
} from "@kensio/rainlytics/cdk";

const logs = new LogBucket(this, "RainlyticsLogs");

const delivery = new CloudFrontLogDelivery(this, "RainlyticsDelivery", {
distributionId: "E1EXAMPLE1234",
logBucket: logs.bucket,
});

new LogTable(this, "RainlyticsTable", { deliveries: [delivery] });
new QueryWorkgroup(this, "RainlyticsQueries");
```

The table projects its partitions, so a query naming a day reads that day and
is billed for those bytes. No crawler runs over the bucket and no partition is
ever registered.

The workgroup bounds what one query may scan. Athena bills per byte and says
nothing at the time, so a query that names no partition is the one mistake here
that costs money quietly. It fails at the point it is run instead.

That stack has to be in us-east-1, which is the only region CloudFront log
delivery can be configured from. See the [log bucket](docs/log-bucket/), [log
delivery](docs/log-delivery/), [log table](docs/log-table/) and [query
workgroup](docs/query-workgroup/) pages.

A `rainlytics` command ships beside them, and answers the questions people
ask most without any SQL:
The command line reads those summaries:

```bash
npx @kensio/rainlytics pageviews --last 7d
rainlytics pageviews --last 7d
```

```text
path views
----------- -----
/ 412
/liju/ 208
/grammar/ 97
/articles/ 208
/pricing/ 97
```

`referrers`, `browsers`, `status-codes` and `cache-hit-ratio` are the others,
and `searches` counts what people typed into a search box. `javascript-errors` and
`web-vitals` read optional browser reports once a deployment opts into those
rollups. `query` takes SQL for anything else. Crawlers are left out by default,
which on the reference site is 39% of the traffic in a typical hour.
Reading a stored answer costs one S3 GET per summary window. Use `--query` when you need Athena to
calculate a fresh answer from the raw logs.

It authenticates through the AWS SDK's default credential chain and writes
JSON, CSV or a table, defaulting to the table at a terminal and to JSON when
it is piped. What a run read and what that came to goes to standard error, so
a pipeline reads rows and a person still sees the price. See the [command
line](docs/command-line/), [rollups](docs/rollups/),
[searches](docs/searches/) and [query](docs/query/) pages.
```bash
rainlytics pageviews --last 2h --query
```

One more construct runs those questions on a schedule and writes each answer
to S3:
Rainlytics has no dashboard. The command line uses the AWS SDK credential chain, including AWS IAM
Identity Center profiles, assumed roles and workload credentials. Output is a table at a terminal
and JSON in a pipe. CSV is available with `--output csv`.

```typescript
import { RollupSummaries } from "@kensio/rainlytics/cdk";
## What Rainlytics reports

new RollupSummaries(this, "RainlyticsSummaries", { table, workgroup });
```
The default scheduled questions cover:

Each question is asked once per hour and once per day, on a lag long enough
for CloudFront to have delivered the window. The named questions above then
read those answers, and a week of pageviews costs 29 GETs and about a
hundredth of a cent. `--query` sends the question to Athena for a fresher
answer, at what a query costs. See the [summary
schedule](docs/summary-schedule/) and [rollup summaries](docs/summaries/)
pages.
- pageviews by path
- referrers by host
- browsers and device classes
- HTTP status codes
- CloudFront cache hit ratio
- search terms from a search page

The same construct precomputes one JSON report for each closed calendar day,
week, month and year. The command reads one without running Athena again:
The same deployment also writes reports for closed days, weeks, months and years.

```bash
rainlytics report month 2026-07 --compare --summaries rainlytics-summaries-1a2b
rainlytics report month 2026-08 --compare
```

The whole versioned document goes to standard output. `--compare` derives changes against the
preceding month from two stored reports. The [calendar reports](docs/reports/) page defines its
periods, sections, comparison rules and S3 keys.

The optional beacon covers what the access log cannot see. A construct answers
a collection path with a 204 at the CloudFront edge, and a module bundled into
the site's own JavaScript reports to it:

```typescript
import { BeaconPath } from "@kensio/rainlytics/cdk";

new BeaconPath(this, "RainlyticsBeacon", { distribution, origin });
```
The browser module can add route changes and custom events. Separate imports collect Core Web
Vitals and uncaught JavaScript errors, so sites only download the features they use.

```typescript
import { startBeacon } from "@kensio/rainlytics/beacon";
import { reportErrors } from "@kensio/rainlytics/beacon/errors";
import { reportVitals } from "@kensio/rainlytics/beacon/vitals";

const beacon = startBeacon();

reportVitals(beacon);
reportErrors(beacon, {
redact: (message) => message.replace(/\S+@\S+/gu, "[email]"),
});

beacon.report({ event: "signup", page: location.pathname });
```

Route changes in a single-page app report themselves. The request stops at the
edge, and CloudFront writes it to the same log objects, the same partitions and
the same table as every page request, so the beacon adds rows rather than a
pipeline. It weighs 586 bytes gzipped, sends no cookies, generates no
identifier, and `pnpm check` fails if it grows past its budget.
The browser sends a GET to `/_rainlytics` on the site's domain. A CloudFront Function returns 204
before the request reaches the cache or origin. CloudFront records the event in the same access log
as every other request.

## AWS resources

Rainlytics ships CDK constructs from `@kensio/rainlytics/cdk`:

Core Web Vitals and uncaught JavaScript errors sit behind imports of their own,
so a site pays for what it asked for:
- `LogBucket` stores the raw CloudFront logs.
- `CloudFrontLogDelivery` configures standard logging v2.
- `LogTable` creates the Glue database and projected table.
- `QueryWorkgroup` adds an Athena scan limit and a results bucket.
- `RollupQueries` saves the generated SQL in Athena.
- `RollupSummaries` schedules rollups and calendar reports.
- `BeaconPath` adds the optional first-party collection path.

The pipeline uses S3, Glue, Athena, Lambda, EventBridge Scheduler and CloudFront. These services are
priced by requests, bytes or execution time. Rainlytics creates no server, provisioned database,
stream or cluster with an hourly capacity charge.

Scheduled Athena queries still have a minimum billed scan, including on a site with no traffic.
See [Summary schedule](docs/summary-schedule/) for the query count and cost model.

## Package entry points

```typescript
import { pageviews, rollups } from "@kensio/rainlytics";
import { LogBucket, RollupSummaries } from "@kensio/rainlytics/cdk";
import { startBeacon } from "@kensio/rainlytics/beacon";
import { reportVitals } from "@kensio/rainlytics/beacon/vitals";
import { reportErrors } from "@kensio/rainlytics/beacon/errors";

reportVitals(beacon);
reportErrors(beacon);
```

That is TTFB, FCP, LCP and CLS, plus what an uncaught error said. All of it
together comes to 1349 bytes gzipped. See the [browser beacon](docs/beacon/),
[beacon path](docs/beacon-path/), [beacon events](docs/beacon-events/),
[JavaScript errors](docs/javascript-errors/) and [Web Vitals](docs/web-vitals/)
pages.

## Status

Experimental and pre-1.0. The construct API changes without a major version
behind it, and the only consumer so far is the maintainer's own sites.
The CDK dependencies are optional peers. Installing Rainlytics for its command line or browser
module does not install `aws-cdk-lib` or `constructs` unless your project requests them.

## Links
## Project status

[rainlytics.com](https://rainlytics.com) is the canonical home.
`rainlytics.dev`, `rainlytics.net` and `rainlytics.app` redirect to it.
Rainlytics is experimental and pre-1.0. Its construct and command interfaces can change without a
major version. The maintainer currently runs it on their own sites.

Rainlytics is written by [Kensio Software](https://kensiosoftware.co.uk) alone.
Sole authorship is what leaves the licence and the direction free to change
later, so pull requests are closed. Issues and bug reports are welcome.
Rainlytics is written by [Kensio Software](https://kensiosoftware.co.uk). Issues and bug reports are
welcome. Pull requests are closed so the project retains a single copyright holder.

## License

Apache 2.0. See [LICENSE](LICENSE).

Rainlytics is an independent open-source project with no affiliation with,
sponsorship from, or endorsement by Amazon or AWS. The name is a nod to
rainforests.
Rainlytics is an independent open-source project. Amazon and AWS do not sponsor, endorse or
affiliate with it. The name is a reference to rainforests.
74 changes: 34 additions & 40 deletions docs/README.md
Original file line number Diff line number Diff line change
@@ -1,42 +1,36 @@
# Rainlytics documentation

Rainlytics is experimental and pre-1.0. The construct API moves without a major version behind it,
because the only consumer so far is the maintainer's own sites.

Pages here are copied to [rainlytics.com](https://rainlytics.com) by that site's scaffold. Each one
needs an H1 and a trailing `<!-- card -->` block. `scripts/sh/docs-check.sh` holds the contract and
runs on every `pnpm check`.

## Constructs

- [Beacon path](beacon-path/), which answers the beacon's collection path with a 204 at the edge.
- [Log bucket](log-bucket/), where CloudFront delivers raw access logs.
- [Log delivery](log-delivery/), which points a distribution at that bucket.
- [Log table](log-table/), the Glue table Athena reads what landed there.
- [Query workgroup](query-workgroup/), which bounds what one query can scan and cost.
- [Rollup queries](rollups/#the-same-sql-saved-in-the-console), the same SQL saved in Athena.
- [Summary schedule](summary-schedule/), which computes the questions on a timer and stores the
answers.

## In the browser

- [Browser beacon](beacon/), which reports the route changes and custom events a server log cannot
see.

## Reading the data back

- [Command line](command-line/), the `rainlytics` command and what it writes.
- [Rollups](rollups/), the named questions, what each counts, and how to write one of your own.
- [Beacon events](beacon-events/), what the beacon reported, with a flood of it bounded.
- [JavaScript errors](javascript-errors/), uncaught exceptions and rejections by page and message.
- [Web Vitals](web-vitals/), p75 for each vital reported through the beacon.
- [Searches](searches/), what readers typed into a search box.
- [Query](query/), running SQL against the log table with `rainlytics query`.
- [Rollup summaries](summaries/), the schema for the precomputed answers the commands read.
- [Calendar reports](reports/), the versioned document for several questions over one closed period.
- [Counting visitors](visitors/), what a visitor count means and over what window.

## Cost

- [Abusing the collection path](abuse/), what an open collection path exposes, and the prices for
containing it.
Start with [Getting started](getting-started/) to deploy Rainlytics and read your first pageview
report. [What Rainlytics is](https://rainlytics.com/guides/what-rainlytics-is/) explains the design
and cost model.

## Set up the pipeline

- [Getting started](getting-started/) covers installation, deployment and the first command.
- [Log bucket](log-bucket/) describes raw log storage and retention.
- [Log delivery](log-delivery/) connects a CloudFront distribution to the bucket.
- [Log table](log-table/) creates the projected Glue table that Athena reads.
- [Query workgroup](query-workgroup/) limits each Athena query and stores its results.
- [Summary schedule](summary-schedule/) precomputes common questions and calendar reports.

## Read analytics

- [Command line](command-line/) covers credentials, regions, output formats and exit codes.
- [Rollups](rollups/) defines the named analytics questions.
- [Searches](searches/) counts terms submitted to search pages.
- [Query](query/) runs ad-hoc SQL through Athena.
- [Rollup summaries](summaries/) documents the stored summary format.
- [Calendar reports](reports/) documents reports for closed calendar periods.
- [Counting visitors](visitors/) explains visitor identity and the required salt.

## Add browser measurements

- [Beacon path](beacon-path/) adds the first-party collection route to CloudFront.
- [Browser beacon](beacon/) reports SPA routes and custom events.
- [Beacon events](beacon-events/) counts custom events and limits repeated identical events.
- [Web Vitals](web-vitals/) reports p75 performance measurements.
- [JavaScript errors](javascript-errors/) groups errors by page and message.
- [Collection-path abuse](abuse/) explains the cost and filtering limits of an open endpoint.

Every topic page lives in its own directory as `README.md`. The website copies these pages into its
Starlight content tree.
Loading