diff --git a/README.md b/README.md index 2989a89..cb4733b 100644 --- a/README.md +++ b/README.md @@ -1,176 +1,130 @@ # Rainlytics -Self-hosted web analytics for AWS sites, built on CloudFront logs. +Rainlytics is self-hosted web analytics for sites served by Amazon CloudFront. -[rainlytics.com](https://rainlytics.com "Rainlytics documentation") +It collects most measurements from CloudFront access logs. Your pages need no analytics +JavaScript, no third-party script tag and no connection to an analytics provider. Raw logs and +derived reports stay in your AWS account. -Rainlytics runs the whole analytics pipeline inside your own AWS account. Most -of what it reports is derived from the CloudFront access logs your distribution -already writes. A measured page downloads no analytics JavaScript, opens no -extra connection, and resolves no extra hostname. +An optional browser module records SPA route changes, custom events, Core Web Vitals and JavaScript +errors. The module goes into your existing bundle and sends events to your own domain. -An optional beacon covers what an access log cannot see accurately, such as -route changes in a single-page app, Core Web Vitals, JavaScript errors and -events the site raises itself. It is bundled into the site's own JavaScript and -reports back through the site's own domain (no second host, no separate script -tag). +[Read the documentation](https://rainlytics.com) or follow [Getting started](docs/getting-started/) +to deploy the first working pipeline. -Everything runs on usage-priced AWS services, batched and precomputed on a -schedule rather than processed per request. Nothing in the pipeline is always -on, and a low-traffic site should cost cents a month. +## How it works -## What works today +CloudFront standard logging v2 writes access logs to S3. Rainlytics creates a Glue table over those +objects and uses partition projection, so there is no Glue crawler. Athena computes common +questions on a schedule. The results are stored as small JSON summaries in S3. -CDK constructs for the collection half of the pipeline. A distribution's access -logs land in an S3 bucket, partitioned and carrying the field set the rollups -will read, and a Glue table describes them for Athena. - -```typescript -import { - CloudFrontLogDelivery, - LogBucket, - LogTable, - QueryWorkgroup, -} from "@kensio/rainlytics/cdk"; - -const logs = new LogBucket(this, "RainlyticsLogs"); - -const delivery = new CloudFrontLogDelivery(this, "RainlyticsDelivery", { - distributionId: "E1EXAMPLE1234", - logBucket: logs.bucket, -}); - -new LogTable(this, "RainlyticsTable", { deliveries: [delivery] }); -new QueryWorkgroup(this, "RainlyticsQueries"); -``` - -The table projects its partitions, so a query naming a day reads that day and -is billed for those bytes. No crawler runs over the bucket and no partition is -ever registered. - -The workgroup bounds what one query may scan. Athena bills per byte and says -nothing at the time, so a query that names no partition is the one mistake here -that costs money quietly. It fails at the point it is run instead. - -That stack has to be in us-east-1, which is the only region CloudFront log -delivery can be configured from. See the [log bucket](docs/log-bucket/), [log -delivery](docs/log-delivery/), [log table](docs/log-table/) and [query -workgroup](docs/query-workgroup/) pages. - -A `rainlytics` command ships beside them, and answers the questions people -ask most without any SQL: +The command line reads those summaries: ```bash -npx @kensio/rainlytics pageviews --last 7d +rainlytics pageviews --last 7d ``` ```text path views ----------- ----- / 412 -/liju/ 208 -/grammar/ 97 +/articles/ 208 +/pricing/ 97 ``` -`referrers`, `browsers`, `status-codes` and `cache-hit-ratio` are the others, -and `searches` counts what people typed into a search box. `javascript-errors` and -`web-vitals` read optional browser reports once a deployment opts into those -rollups. `query` takes SQL for anything else. Crawlers are left out by default, -which on the reference site is 39% of the traffic in a typical hour. +Reading a stored answer costs one S3 GET per summary window. Use `--query` when you need Athena to +calculate a fresh answer from the raw logs. -It authenticates through the AWS SDK's default credential chain and writes -JSON, CSV or a table, defaulting to the table at a terminal and to JSON when -it is piped. What a run read and what that came to goes to standard error, so -a pipeline reads rows and a person still sees the price. See the [command -line](docs/command-line/), [rollups](docs/rollups/), -[searches](docs/searches/) and [query](docs/query/) pages. +```bash +rainlytics pageviews --last 2h --query +``` -One more construct runs those questions on a schedule and writes each answer -to S3: +Rainlytics has no dashboard. The command line uses the AWS SDK credential chain, including AWS IAM +Identity Center profiles, assumed roles and workload credentials. Output is a table at a terminal +and JSON in a pipe. CSV is available with `--output csv`. -```typescript -import { RollupSummaries } from "@kensio/rainlytics/cdk"; +## What Rainlytics reports -new RollupSummaries(this, "RainlyticsSummaries", { table, workgroup }); -``` +The default scheduled questions cover: -Each question is asked once per hour and once per day, on a lag long enough -for CloudFront to have delivered the window. The named questions above then -read those answers, and a week of pageviews costs 29 GETs and about a -hundredth of a cent. `--query` sends the question to Athena for a fresher -answer, at what a query costs. See the [summary -schedule](docs/summary-schedule/) and [rollup summaries](docs/summaries/) -pages. +- pageviews by path +- referrers by host +- browsers and device classes +- HTTP status codes +- CloudFront cache hit ratio +- search terms from a search page -The same construct precomputes one JSON report for each closed calendar day, -week, month and year. The command reads one without running Athena again: +The same deployment also writes reports for closed days, weeks, months and years. ```bash -rainlytics report month 2026-07 --compare --summaries rainlytics-summaries-1a2b +rainlytics report month 2026-08 --compare ``` -The whole versioned document goes to standard output. `--compare` derives changes against the -preceding month from two stored reports. The [calendar reports](docs/reports/) page defines its -periods, sections, comparison rules and S3 keys. - -The optional beacon covers what the access log cannot see. A construct answers -a collection path with a 204 at the CloudFront edge, and a module bundled into -the site's own JavaScript reports to it: - -```typescript -import { BeaconPath } from "@kensio/rainlytics/cdk"; - -new BeaconPath(this, "RainlyticsBeacon", { distribution, origin }); -``` +The browser module can add route changes and custom events. Separate imports collect Core Web +Vitals and uncaught JavaScript errors, so sites only download the features they use. ```typescript import { startBeacon } from "@kensio/rainlytics/beacon"; +import { reportErrors } from "@kensio/rainlytics/beacon/errors"; +import { reportVitals } from "@kensio/rainlytics/beacon/vitals"; const beacon = startBeacon(); + +reportVitals(beacon); +reportErrors(beacon, { + redact: (message) => message.replace(/\S+@\S+/gu, "[email]"), +}); + beacon.report({ event: "signup", page: location.pathname }); ``` -Route changes in a single-page app report themselves. The request stops at the -edge, and CloudFront writes it to the same log objects, the same partitions and -the same table as every page request, so the beacon adds rows rather than a -pipeline. It weighs 586 bytes gzipped, sends no cookies, generates no -identifier, and `pnpm check` fails if it grows past its budget. +The browser sends a GET to `/_rainlytics` on the site's domain. A CloudFront Function returns 204 +before the request reaches the cache or origin. CloudFront records the event in the same access log +as every other request. + +## AWS resources + +Rainlytics ships CDK constructs from `@kensio/rainlytics/cdk`: -Core Web Vitals and uncaught JavaScript errors sit behind imports of their own, -so a site pays for what it asked for: +- `LogBucket` stores the raw CloudFront logs. +- `CloudFrontLogDelivery` configures standard logging v2. +- `LogTable` creates the Glue database and projected table. +- `QueryWorkgroup` adds an Athena scan limit and a results bucket. +- `RollupQueries` saves the generated SQL in Athena. +- `RollupSummaries` schedules rollups and calendar reports. +- `BeaconPath` adds the optional first-party collection path. + +The pipeline uses S3, Glue, Athena, Lambda, EventBridge Scheduler and CloudFront. These services are +priced by requests, bytes or execution time. Rainlytics creates no server, provisioned database, +stream or cluster with an hourly capacity charge. + +Scheduled Athena queries still have a minimum billed scan, including on a site with no traffic. +See [Summary schedule](docs/summary-schedule/) for the query count and cost model. + +## Package entry points ```typescript +import { pageviews, rollups } from "@kensio/rainlytics"; +import { LogBucket, RollupSummaries } from "@kensio/rainlytics/cdk"; +import { startBeacon } from "@kensio/rainlytics/beacon"; import { reportVitals } from "@kensio/rainlytics/beacon/vitals"; import { reportErrors } from "@kensio/rainlytics/beacon/errors"; - -reportVitals(beacon); -reportErrors(beacon); ``` -That is TTFB, FCP, LCP and CLS, plus what an uncaught error said. All of it -together comes to 1349 bytes gzipped. See the [browser beacon](docs/beacon/), -[beacon path](docs/beacon-path/), [beacon events](docs/beacon-events/), -[JavaScript errors](docs/javascript-errors/) and [Web Vitals](docs/web-vitals/) -pages. - -## Status - -Experimental and pre-1.0. The construct API changes without a major version -behind it, and the only consumer so far is the maintainer's own sites. +The CDK dependencies are optional peers. Installing Rainlytics for its command line or browser +module does not install `aws-cdk-lib` or `constructs` unless your project requests them. -## Links +## Project status -[rainlytics.com](https://rainlytics.com) is the canonical home. -`rainlytics.dev`, `rainlytics.net` and `rainlytics.app` redirect to it. +Rainlytics is experimental and pre-1.0. Its construct and command interfaces can change without a +major version. The maintainer currently runs it on their own sites. -Rainlytics is written by [Kensio Software](https://kensiosoftware.co.uk) alone. -Sole authorship is what leaves the licence and the direction free to change -later, so pull requests are closed. Issues and bug reports are welcome. +Rainlytics is written by [Kensio Software](https://kensiosoftware.co.uk). Issues and bug reports are +welcome. Pull requests are closed so the project retains a single copyright holder. ## License Apache 2.0. See [LICENSE](LICENSE). -Rainlytics is an independent open-source project with no affiliation with, -sponsorship from, or endorsement by Amazon or AWS. The name is a nod to -rainforests. +Rainlytics is an independent open-source project. Amazon and AWS do not sponsor, endorse or +affiliate with it. The name is a reference to rainforests. diff --git a/docs/README.md b/docs/README.md index 1fdbf89..55d9f7f 100644 --- a/docs/README.md +++ b/docs/README.md @@ -1,42 +1,36 @@ # Rainlytics documentation -Rainlytics is experimental and pre-1.0. The construct API moves without a major version behind it, -because the only consumer so far is the maintainer's own sites. - -Pages here are copied to [rainlytics.com](https://rainlytics.com) by that site's scaffold. Each one -needs an H1 and a trailing `` block. `scripts/sh/docs-check.sh` holds the contract and -runs on every `pnpm check`. - -## Constructs - -- [Beacon path](beacon-path/), which answers the beacon's collection path with a 204 at the edge. -- [Log bucket](log-bucket/), where CloudFront delivers raw access logs. -- [Log delivery](log-delivery/), which points a distribution at that bucket. -- [Log table](log-table/), the Glue table Athena reads what landed there. -- [Query workgroup](query-workgroup/), which bounds what one query can scan and cost. -- [Rollup queries](rollups/#the-same-sql-saved-in-the-console), the same SQL saved in Athena. -- [Summary schedule](summary-schedule/), which computes the questions on a timer and stores the - answers. - -## In the browser - -- [Browser beacon](beacon/), which reports the route changes and custom events a server log cannot - see. - -## Reading the data back - -- [Command line](command-line/), the `rainlytics` command and what it writes. -- [Rollups](rollups/), the named questions, what each counts, and how to write one of your own. -- [Beacon events](beacon-events/), what the beacon reported, with a flood of it bounded. -- [JavaScript errors](javascript-errors/), uncaught exceptions and rejections by page and message. -- [Web Vitals](web-vitals/), p75 for each vital reported through the beacon. -- [Searches](searches/), what readers typed into a search box. -- [Query](query/), running SQL against the log table with `rainlytics query`. -- [Rollup summaries](summaries/), the schema for the precomputed answers the commands read. -- [Calendar reports](reports/), the versioned document for several questions over one closed period. -- [Counting visitors](visitors/), what a visitor count means and over what window. - -## Cost - -- [Abusing the collection path](abuse/), what an open collection path exposes, and the prices for - containing it. +Start with [Getting started](getting-started/) to deploy Rainlytics and read your first pageview +report. [What Rainlytics is](https://rainlytics.com/guides/what-rainlytics-is/) explains the design +and cost model. + +## Set up the pipeline + +- [Getting started](getting-started/) covers installation, deployment and the first command. +- [Log bucket](log-bucket/) describes raw log storage and retention. +- [Log delivery](log-delivery/) connects a CloudFront distribution to the bucket. +- [Log table](log-table/) creates the projected Glue table that Athena reads. +- [Query workgroup](query-workgroup/) limits each Athena query and stores its results. +- [Summary schedule](summary-schedule/) precomputes common questions and calendar reports. + +## Read analytics + +- [Command line](command-line/) covers credentials, regions, output formats and exit codes. +- [Rollups](rollups/) defines the named analytics questions. +- [Searches](searches/) counts terms submitted to search pages. +- [Query](query/) runs ad-hoc SQL through Athena. +- [Rollup summaries](summaries/) documents the stored summary format. +- [Calendar reports](reports/) documents reports for closed calendar periods. +- [Counting visitors](visitors/) explains visitor identity and the required salt. + +## Add browser measurements + +- [Beacon path](beacon-path/) adds the first-party collection route to CloudFront. +- [Browser beacon](beacon/) reports SPA routes and custom events. +- [Beacon events](beacon-events/) counts custom events and limits repeated identical events. +- [Web Vitals](web-vitals/) reports p75 performance measurements. +- [JavaScript errors](javascript-errors/) groups errors by page and message. +- [Collection-path abuse](abuse/) explains the cost and filtering limits of an open endpoint. + +Every topic page lives in its own directory as `README.md`. The website copies these pages into its +Starlight content tree. diff --git a/docs/abuse/README.md b/docs/abuse/README.md index 1e5fe0d..5c7d69d 100644 --- a/docs/abuse/README.md +++ b/docs/abuse/README.md @@ -1,133 +1,65 @@ -# What abusing the collection path costs +# Collection-path abuse -The beacon reports to a path on the site's own domain, and CloudFront records the request in the -access log like any other. The path is open and unauthenticated. Anybody can send that URL a -million times and have every one of them counted, carrying a page value naming a page nobody opened and an event that -never happened. +The browser collection path is public. Any client can send requests to it and invent event names, +pages and values. -Two things follow, and they want different answers. The counts recover. The money is spent. +This risk also exists for access-log pageviews. A client can request a real page repeatedly and +create valid-looking rows. Server logs record requests. A request alone provides no proof that a +person read the response. -## Layer 1 is open in the same way +## Protect the reported count -This comes first because the beacon looks like the thing that opened the door. +The `beaconEvents` rollup caps one visitor's identical events at 60 per hour. The standard bot +filter also removes clients that identify themselves with common crawler names. -A site's own pages take a request from anybody. A million requests for a real page put a million -rows in the log. Each one is a GET that answered HTML and succeeded, and that is the whole of what -`pageviews` asks of a row, so the count follows the flood up. The [crawler -filter](../rollups/#crawlers-are-most-of-the-traffic) catches a flood naming itself a bot and -nothing else about it. Every analytics product built on server logs works this way. A log records -what arrived and has no way to ask why. +These rules protect the derived count. They do not remove requests from the raw log. You can change +a query and recompute a poisoned window while its raw objects still exist. -What layer 2 adds is a forged page value and events nobody caused. The gap is narrower than it -looks. A spammed page request already lies about which page was read, and it transfers the page body -to do it. A spammed beacon request carries no body in either direction. +A client can avoid the cap by rotating addresses, user agents, pages or event names. Treat event +counts as signals from an open endpoint. -## The counts recover +## Request cost is final -The raw store is immutable and every rollup is rebuilt from it. A poisoned window is a re-run under -a better filter. +Every abusive request can incur: -[#104](https://github.com/KensioSoftware/rainlytics/issues/104) chose that filter and -[`beacon-events`](../beacon-events/) applies it. One visitor's identical events are counted no more -than 60 times an hour, which is one a minute from one person, on one page, of one event name. It -sits in the rollup query, beside the crawler filter every question already applies. The raw store -keeps every row and the query decides what to count. A rule that turns out to be wrong is another -re-run. +- a CloudFront request and CloudFront Function invocation +- S3 request and storage cost for the delivered log record +- Athena scan cost whenever a query reads the affected partition -The [log bucket's](../log-bucket/) expiry is the outer limit on this. A window that has aged past it -has no rows left to recount, under any filter at all. A year is the default. +Filtering later changes the report only. The Athena workgroup limits one query's scan. CloudFront +and S3 charges remain outside that limit. -## The money is spent +Every component is usage-priced. A large request flood therefore creates a large variable bill even +though the normal deployment has no fixed monthly capacity. -A re-run fixes a number. Nothing re-runs a bill. Every spammed request buys two charges outright -and arms a third, and no filter written afterwards takes any of them back. +## Add WAF when its fixed cost is justified -**A CloudFront request.** The distribution charges per request at its own rate, and that charge -lands on the CDN bill whether Rainlytics is installed or not. A request for a real page costs the -same and transfers a page body on top of it. Whatever answers the collection path is priced per hit -too, and [#99](https://github.com/KensioSoftware/rainlytics/issues/99) settled on a [CloudFront -Function](../beacon-path/) at $0.10 per million invocations. +AWS WAF can keep request counts at the edge and apply a rate-based rule to the collection path. A +new web ACL has a monthly charge, each rule has another monthly charge, and request inspection is +also billed. -**A log record, kept for the bucket's retention.** -[#9](https://github.com/KensioSoftware/rainlytics/issues/9) measured the log store at $0.084 a month -on a site serving 137,000 requests a day, which works out near $0.02 per million requests. It splits -between one PUT per delivered object and steady-state storage under the 370-day expiry. CloudFront -delivers into the bucket at no charge, which made that figure the whole of what Rainlytics itself -cost on that site. A flood pays the rate on the way in and then pays the storage every month until -the expiry drops it. +At the standard published rates used by the project, one web ACL and one rate-based rule begin at +$6 per month before request charges. This is much larger than the normal log-storage cost of a quiet +site. Rainlytics therefore leaves WAF configuration to the site. -**Bytes that a query over the window scans.** This is the armed one. Athena bills $5.00 per terabyte, -and the charge arrives only when something reads the window (a scheduled rollup, or a `--query` run -for a fresher answer). Spammed rows sit in the same objects as real ones and no partition predicate -tells them apart, so each run that covers the window reads them again for as long as it stays in -range. +WAF is cheaper to add when the distribution already has a web ACL. Define the rule in the site's +own CDK app and scope it to `defaultBeaconPath` or the custom path passed to `BeaconPath`. -That third charge already has a ceiling. The [query workgroup's](../query-workgroup/) bytes-scanned -cutoff fails a query at ten gibibytes, which caps one query near five cents whatever the flood put in -the window. It binds queries naming the workgroup, being `rainlytics` unless a deployment renamed it. -Athena's own `primary` workgroup has no cutoff, and a query landing there is uncapped. The first two -charges have no ceiling anywhere. +CloudFront Functions start each request without writable state, so they cannot implement a counter. +CloudFront KeyValueStore is read-only from function code. AWS Shield Standard protects the network +layer. Application-level rate limiting requires WAF. -## AWS WAF, and why it stays out of the default +## Monitor spend -WAF is the one place at the edge where a request count can be kept, and it is priced in the open. -Read from the AWS WAF pricing page on 2026-08-29: +Create an AWS Budget for the account or workload and alert above its normal monthly range. The first +two AWS Budgets in an account have no charge under the standard pricing described by AWS. -- $5.00 a month per web ACL -- $1.00 a month per rule -- $0.60 per million requests inspected - -A rate-based rule is an ordinary rule at $1.00. So the smallest configuration that would help, one -web ACL carrying one rate-based rule on the collection path, is $6.00 a month before a single request -reaches it. - -Set that beside the $0.084 a month #9 measured. WAF is a fixed floor around seventy times the log -store it protects, and it is billed in full in a quiet month when nobody attacks anything. Every -other charge on this page is priced by use, and this would be the largest line on a quiet site's -bill. - -That answer flips for a site already running a web ACL for other reasons. The $5.00 is paid, the -rule is $1.00, and the collection path joins something that exists. The default is for a site -installing Rainlytics, where the ACL would exist for this alone. - -## Why the count has to live in WAF - -Rate limiting needs a count that survives between requests, and the edge has nowhere to keep one. - -- **CloudFront Functions** hold no state between invocations. A function sees one request and - forgets it. -- **CloudFront KeyValueStore** is read-only from function code. A function reads what a deploy put - there and cannot write a counter back. -- **Shield Standard** comes at no charge and works at the network layer. Ten well-formed HTTPS - requests a second look like traffic to it. -- **Shield Advanced** carries the application-layer protection and costs $3,000 a month. - -## A budget alarm is the honest answer - -An exposure that outlasts every attempt to prevent it is one to be told about. AWS Budgets gives an -account its first two budgets at no charge, and a cost alarm is one of them. - -Put one on the account carrying the distribution and the log bucket, with a threshold above what a -quiet month costs (#9's figure is the right shape for a site of that size, and a month of real -billing is better). An alert firing at twice a normal month is a flood in progress. The decision -about WAF is then taken with a bill in hand, which beats guessing at one during a deploy. - -## No WAF construct ships here - -Every resource Rainlytics creates is priced by use, and a construct putting $6.00 a month into the -default path would break that for every site installing it. Whether the $6.00 is worth paying -depends on what a site is worth attacking, what else its account already runs, and what its owner -wants to spend. That is the site's decision, and it is taken with information the library lacks. - -A site taking it writes the web ACL in its own CDK app and associates it with the distribution. The -collection path is `/_rainlytics` unless a site names another, and it is exported as -`defaultBeaconPath` from the package root, so a rate-based rule can scope itself to the same path -the beacon reports to. +A budget alert detects CloudFront, function, storage and query growth together. Use the resulting +traffic and cost data to decide whether a WAF rule is worth its monthly floor. diff --git a/docs/beacon-events/README.md b/docs/beacon-events/README.md index 4c63499..564f496 100644 --- a/docs/beacon-events/README.md +++ b/docs/beacon-events/README.md @@ -1,150 +1,90 @@ # Beacon events -Counts what the beacon reported, by the page an event happened on and the name it was reported -under, with a flood of identical events bounded. +The `beaconEvents` rollup counts browser events by page and event name. ```typescript import { beaconEvents, defaultBeaconPath, rollups } from "@kensio/rainlytics"; -const beaconPath = { paths: [defaultBeaconPath] }; +const questions = [...rollups, beaconEvents]; +const requests = { + "beacon-events": { paths: [defaultBeaconPath] }, +}; -new RollupQueries(this, "RainlyticsQueries", { +new RollupQueries(this, "SavedQueries", { table, workgroup, - rollups: [...rollups, beaconEvents], - requests: { "beacon-events": beaconPath }, + rollups: questions, + requests, }); -``` - -```bash -rainlytics saved-query beacon-events -``` - -```text -page event events ----------- ------ ------ -/liju/ route 412 -/liju/ vital 408 -/grammar/ route 97 -``` -The page comes out of the query string rather than out of the request, because the request was sent -to the beacon's own path. A route change in a single-page app is what that exists for, where the -address bar has moved and no request was made. - -## A site opts into it - -A deployment using the default rollup list computes its six questions. This one waits to be asked -for, so a site with no beacon leaves it out and computes none of it. Layer 2 is optional, and a -scheduled question over rows nobody writes is an Athena charge per window for an empty answer. - -A site running the beacon adds it to both constructs, once to save the query and once to compute it -on a schedule: - -```typescript -new RollupSummaries(this, "RainlyticsSummaries", { +new RollupSummaries(this, "Summaries", { table, workgroup, - rollups: [...rollups, beaconEvents], - requests: { "beacon-events": beaconPath }, + rollups: questions, + requests, }); ``` -That is 50 more Athena queries a day on the two cadences, which comes to about 8 cents a month. -`rainlytics saved-query beacon-events` runs the saved copy through Athena for a fresh answer, and -the summaries it writes are JSON on S3 in the [shape every summary takes](../summaries/). - -## Narrow it to the collection path - -`--path` is what separates a beacon event from every other query string in the same log. `?v=3` on a -stylesheet is an ordinary thing for a site to serve, and this question reads a `v` parameter to find -its envelope. - -The `requests` entry above is that narrowing for a saved copy and for a schedule. A site that moved -its beacon names its own path there, in the one place, and the [beacon path](../beacon-path/) -construct takes the same value. +Run the saved query for a fresh answer: -## One visitor, sixty events an hour - -The collection path is open and unauthenticated by design. Anybody can send its URL a million times -and have every one of them counted, under a page value naming a page nobody opened. - -So this question counts one visitor's identical events no more than 60 times an hour. That is one a -minute from one person, on one page, of one event name. A reader who moves around a site produces -events on several pages and is capped on each of them separately. A client sending the same URL a -million times contributes 60. - -The hour is the row's own, taken from its timestamp rather than from the window being computed. An -hourly summary and the daily summary covering it therefore apply the same cap, and the 24 hours of a -day add up to the day. - -Two rules stack here. The [crawler filter](../rollups/#crawlers-are-most-of-the-traffic) every -question applies has already taken a flood that names itself a bot. The cap is for the quiet -kind. - -### Why a cap per visitor - -Two other rules were considered, and both fall to what the beacon is for. -[#104](https://github.com/KensioSoftware/rainlytics/issues/104) has them. +```bash +rainlytics saved-query beacon-events +``` -**Dropping events whose page never appears as a pageview in the same window.** A route change in a -single-page app has no document request behind it. This rule drops exactly the events layer 2 was -built to collect. +```text +page event events +---------- ------- ------ +/articles/ route 412 +/checkout/ signup 38 +``` -**A cap per path.** A popular page legitimately carries many events. A cap low enough to bound a -flood clips real traffic, and one high enough to spare real traffic lets a flood run underneath it. +The page comes from the event's `p` parameter. The request itself always goes to the collection +path. -A cap per visitor is the one that scales with real popularity. Ten thousand readers count as ten -thousand, and one client counts as one whatever it sends. +## Opt in -### Where it runs out +Beacon events are an optional rollup. An access-log-only deployment produces no event rows and +avoids the empty scheduled queries. -A client rotating addresses counts as many visitors and gets the cap each time. An access log reads -that as a crowd, and the same limit applies to the [visitor count](../visitors/). +Adding the rollup under both default granularities and two-window recomputation adds 50 Athena +queries a day. At Athena's minimum scan and standard rate, this is about eight cents a month before +traffic raises the scan above the minimum. -Every spammed request also costs money before any of this runs, and no filter written afterwards -takes it back. [What abuse of the collection path costs](../abuse/) has the three charges and the -budget alarm that is the honest answer to them. +The request must name the collection path. If `BeaconPath` uses `/_measure`, use the same path in +`requests`. -## It needs the viewer's address +## Repeated-event cap -A visitor here is the address and the user agent CloudFront logged, which is the pair a [visitor -count](../visitors/) is hashed from. Both values stay inside the query. The inner `SELECT` groups by -them and the outer one adds up what that produced, so no address reaches a summary, a result object -or a reader. +The collection path is open and unauthenticated. `beaconEvents` counts one visitor's identical +events at most 60 times an hour. The key contains the visitor, page, event name and log hour. -A deployment delivering [`logFieldNamesWithoutAddress`](../log-delivery/#the-field-set-holds-the-viewers-address) -has no column to key the cap on, and `RollupSummaries` refuses the question at synthesis: +This cap allows real traffic to grow with the audience while limiting one client that repeats the +same event URL. A client can bypass the cap by rotating addresses, user agents, pages or event +names. -```text -beacon-events bounds a flood by capping what one visitor sent, and this deployment's -delivery leaves out c-ip. Either add c-ip to the delivered field set, or leave the -question out of this deployment. Counting beacon events with no cap would report a -flood of a million as a million. -``` +The cap requires the viewer address and user agent from the access log. A delivery using +`logFieldNamesWithoutAddress` cannot schedule `beaconEvents`. Rainlytics rejects that combination +during synthesis. -The cap has no off switch, which is what separates this from the visitor count. A summary without a -visitor count is the same question with one fewer number beside it. A beacon rollup without the cap -is a different question, and it would report a flood of a million as a million. +No address or user agent reaches the summary. The query groups by them internally and writes the +final counts only. -## The raw store keeps everything +## Raw events remain available -The cap is applied by the query and changes no object in the bucket. Every request the beacon path -answered is still a row, spammed ones included, and `rainlytics query` counts them all for anybody -checking what arrived: +The cap changes the query result, not the raw log. Every request remains in S3 until the log bucket +expires it. Use ad-hoc SQL to inspect the full request count: ```sql -SELECT count(*) FROM "rainlytics"."cloudfront_logs" - WHERE year = '2026' AND month = '08' AND day = '23' - AND strpos(cs_uri_stem, '/_rainlytics') = 1 +SELECT count(*) +FROM rainlytics.cloudfront_logs +WHERE year = '2026' AND month = '09' AND day = '01' + AND cs_uri_stem = '/_rainlytics' ``` -A rule that turns out to be wrong is a re-run over rows that are all still there. The [log -bucket's](../log-bucket/) expiry is the outer limit on that, and a year is the default. +See [Collection-path abuse](../abuse/) for request costs and limits. diff --git a/docs/beacon-path/README.md b/docs/beacon-path/README.md index 851380b..74bb17e 100644 --- a/docs/beacon-path/README.md +++ b/docs/beacon-path/README.md @@ -1,194 +1,101 @@ # Beacon path -Answers the beacon's collection path with a 204, at the edge. The construct adds a cache behaviour -to a distribution you already own and attaches a CloudFront Function to it. The request stops at -the edge, ahead of the cache and ahead of the origin, and CloudFront records it in the access log -like every other request. That record is the event. +`BeaconPath` adds a first-party event collection route to an existing CloudFront distribution. ```typescript import { BeaconPath } from "@kensio/rainlytics/cdk"; -new BeaconPath(this, "RainlyticsBeacon", { distribution, origin }); -``` - -`distribution` is the one already serving the site. `origin` is whatever the rest of the site comes -from. Every CloudFront cache behaviour names an origin and no request reaches this one, so pass the -site's own. - -The path defaults to `/_rainlytics`. - -## Plain HTTP is refused - -The behaviour is HTTPS-only by default. A plain HTTP request gets 403 rather than the 204 an HTTPS -request gets. - -`redirect-to-https` would make CloudFront answer an HTTP request with 301 and wait for the browser -to send it again over HTTPS. The beacon's most fragile send is made with `keepalive` while its page -is going away. Making that send depend on a second request creates another point where the event can -be lost. HTTPS-only rejects a request that should not have started on HTTP and needs no follow-up. - -Pass `viewerProtocolPolicy` where a site has a reason to differ: - -```typescript -import { ViewerProtocolPolicy } from "aws-cdk-lib/aws-cloudfront"; - -new BeaconPath(this, "RainlyticsBeacon", { +new BeaconPath(this, "BeaconPath", { distribution, origin, - viewerProtocolPolicy: ViewerProtocolPolicy.REDIRECT_TO_HTTPS, }); ``` -`ViewerProtocolPolicy.ALLOW_ALL` is available too. It deliberately makes the collection path answer -plain HTTP, even if every other behaviour on the distribution redirects or refuses it. +The default route is `/_rainlytics`. A CloudFront Function returns 204 during viewer request, before +the cache or origin. CloudFront still records the request in its access log. + +`distribution` must be the CDK `Distribution` that serves the measured site. Pass any existing site +origin. CloudFront requires an origin on every behavior, but a beacon request never reaches it. -## How an event travels +## Request flow -The beacon puts its payload in the query string and sends a GET: +The browser sends an event in the query string: ```text -GET /_rainlytics?v=1&e=route&p=%2Fliju%2F +GET /_rainlytics?v=1&e=route&p=%2Farticles%2F ``` -CloudFront delivers `cs-uri-query` whatever the cache key and the origin forwarding are set to. The -payload lands in the same log objects, the same partitions and the same Glue table as every page -request. Layer 2 is more rows in the dataset layer 1 already writes, and that is what makes the -beacon nearly free. `src/beacon-events.ts` holds the envelope, and the [log table](../log-table/) -page has the columns. - -The function reads none of it. It returns the same 204 to every request the behaviour matches, and -the payload travels past it into the log. - -## The choice between a function and a cached object - -An empty path can be answered two ways. A CloudFront Function on viewer-request returns a synthetic -204, and a small object on the origin is served from the cache. Both carry the same CloudFront -request charge and the same log delivery, and the invocation charge below is the only difference -between them. [#99](https://github.com/KensioSoftware/rainlytics/issues/99) took the decision on -three things beside cost. - -**Rainlytics can ship the function on its own.** The construct attaches a behaviour and a function -to a distribution you already own, whatever that distribution serves. A cached object needs -something to put a file at the path, which is either the site's build cooperating or Rainlytics -writing into an origin bucket it has been granted. A distribution in front of an ALB, an API or a -third party takes the function and has nowhere to put the object. - -**A flood stays inside CloudFront.** The function answers at the edge and the origin never hears -about it. A cached object serves from cache until its TTL lapses, and the misses reach the origin, -bounded by points of presence times TTL. The collection path is unauthenticated by design. That is -the difference between an abusive client costing CloudFront requests and one arriving at the site -itself. - -**The cache hit ratio stays honest.** `cache-hit-ratio` counts `Hit`, `RefreshHit` and `Miss` -alone. A cached object would count as a `Hit` on nearly every beacon request and lift the ratio for -the whole site. A generated response is none of the three. - -## What it costs +CloudFront writes `cs-uri-query` independently of the cache key and origin forwarding settings. The +function ignores the payload and returns the same empty response for every matching request. -CloudFront Functions are priced per invocation, at $0.10 per million. A viewer-request function runs -before the cache on every request the behaviour matches, so a million beacon events is a million -invocations and ten pence. +The event enters the same S3 objects, Glue table and Athena queries as normal page requests. There +is no separate ingestion API. -Beside that sit the CloudFront request charge and the log delivery, which a million page requests -pay anyway and which the cached object would pay too. Every charge here is per request, and a -quiet site pays for the requests it gets. Prices read from the CloudFront pricing page on -2026-08-29. +## Choose the path -## The cache key leaves the query string out +Reserve a path for the beacon: -The behaviour takes the managed `CachingOptimized` policy, which keys on the path alone. - -A viewer-request function that returns a response ends the request before the cache is consulted, -which leaves this path with no cache entry to hold. The policy still matters. It is what the path -falls back to if the function is ever removed or fails to associate, and a policy carrying the query -string would make every event its own cache key and send every one of them to the origin. Keying on -the path collapses that same flood into a single entry. - -Pass `cachePolicy` for a site that standardises on one of its own. Whatever it is, keep the query -string out of the key. - -## How much fits in the URL +```typescript +new BeaconPath(this, "BeaconPath", { + distribution, + origin, + path: "/_measure", +}); +``` -CloudFront accepts 8,192 bytes of path and query string together, and 32,768 bytes for the whole -request including its headers. Above either limit it answers 414 and the event is lost. +Pass the same value to `startBeacon` and to any beacon rollup request. Rainlytics rejects a path +without a leading slash or a path containing a query string. -That leaves several kilobytes for a payload, against the 800 bytes `cf.logCustomData()` allows. It -is why the beacon carries its data in the URL and why no API Gateway and Lambda ingest endpoint is -in the design. Reach for `logCustomData` only where the value has to come from the function itself. +Do not use a real page path. Each event would then look like a request for that page and could +download its body if the edge function were missing. -## Paths CloudFront would never match +## Protocol and cache behavior -A CloudFront path pattern may start with anything, and `*.jpg` is a normal one. A beacon path -written without its leading slash therefore deploys green and matches no request the beacon sends, -and the first sign of it is a dataset with no beacon rows in it. +The path is HTTPS-only by default. Plain HTTP receives 403. A redirect would require a second +request, which is unreliable when the browser sends the event while leaving a page. -The construct refuses that at synthesis, naming the path. It refuses one carrying a query string for -the same reason, since a pattern is matched against the path alone. +Change the policy only when the site requires it: ```typescript -new BeaconPath(this, "RainlyticsBeacon", { +import { ViewerProtocolPolicy } from "aws-cdk-lib/aws-cloudfront"; + +new BeaconPath(this, "BeaconPath", { distribution, origin, - path: "/_collect", + viewerProtocolPolicy: ViewerProtocolPolicy.REDIRECT_TO_HTTPS, }); ``` -Give it a path the site keeps free. Pointing the beacon at a page would count every event as a -view of that page, and download the page body a second time. - -## Browsers keep no copy of the answer - -The function sends `cache-control: no-store`. The same event on the same page produces the same URL -twice, and a browser holding a cached 204 would answer the second one out of its own cache. That -event would reach no log. +The default managed cache policy excludes query strings from the cache key. The function normally +ends the request before the cache, but the safe fallback is one cached path rather than one key per +event. -## What a beacon row looks like in the log +The response includes `cache-control: no-store`, which prevents a browser from satisfying a repeated +event from its own cache. -Measured on 2026-08-30, from a verification request to the collection path on a deployed -distribution: - -```json -{ - "cs-uri-stem": "/_rainlytics", - "cs-uri-query": "v=1&e=verify119&p=%252Fverify%252F", - "sc-status": "204", - "sc-content-type": "-", - "x-edge-result-type": "FunctionGeneratedResponse" -} -``` +## Limits and cost -**`x-edge-result-type` is `FunctionGeneratedResponse`**, which the CloudFront documentation does not -enumerate. The published list runs `Hit`, `RefreshHit`, `Miss`, `LimitExceeded`, `CapacityExceeded`, -`Error`, `Redirect` and `LambdaExecutionError`, and this is none of them. The same value arrives at -the browser as `x-cache: FunctionGeneratedResponse from cloudfront`. +CloudFront accepts up to 8,192 bytes for the path and query string and 32,768 bytes for the complete +request. Events above either limit receive 414, and CloudFront drops their payload. -That is the value [cache hit ratio](../rollups/) needed. It counts a Hit, a RefreshHit and a Miss -and nothing else, so a beacon row falls outside the ratio without the question having to know the -beacon exists. Had it come back as a Hit, every event would have inflated the ratio. -[#119](https://github.com/KensioSoftware/rainlytics/issues/119) is where that was settled. +A viewer-request CloudFront Function costs $0.10 per million invocations at the documented standard +rate. CloudFront request and log storage charges also apply. Every charge scales with requests. -**The payload is encoded twice.** The browser encodes what it sends and CloudFront encodes the -record again, so a `%2F` on the way out is `%252F` in the row. `decodedParameter` in -`src/log-encoding.ts` is the two passes that read it back, and the row above is what they were -written for. +A cached origin object would avoid the function invocation charge, but it would require every site +origin to serve the object, allow occasional origin misses and inflate the cache hit ratio. The +function keeps all event handling at the edge. -Beacon rows land in the status-code rollup as 204s, and `status-codes` takes them back out by path. -Anybody asking that question means to count page requests. [Beacon -events](../beacon-events/) has the rollup that counts the beacon's own rows instead, with the cap -that bounds a flood of them. +## Logged result type -## Permissions for a scoped deploy role - -Untested. An account still on the `AdministratorAccess` that `cdk bootstrap` gives the -CloudFormation execution role deploys this with no IAM work, and the reference site is one. A role -narrowed with `--cloudformation-execution-policies` is expected to need `cloudfront:CreateFunction`, -`cloudfront:DescribeFunction`, `cloudfront:PublishFunction` and `cloudfront:DeleteFunction` for the -function, plus `cloudfront:GetDistribution` and `cloudfront:UpdateDistribution` for the behaviour. -Treat that as a starting point for reading a denial. The [log delivery](../log-delivery/) page has -what happened when the same list was reasoned about rather than deployed. +CloudFront records a successful event with status 204 and result type +`FunctionGeneratedResponse`. The cache hit ratio counts `Hit`, `RefreshHit` and `Miss` only. +`FunctionGeneratedResponse` therefore stays outside the ratio. diff --git a/docs/beacon/README.md b/docs/beacon/README.md index 7f67e9b..a999ce2 100644 --- a/docs/beacon/README.md +++ b/docs/beacon/README.md @@ -1,9 +1,7 @@ # Browser beacon -Reports what a server log cannot see. A single-page app changing route moves the address bar and -makes no request, so CloudFront records nothing and layer 1 counts nothing. The beacon fills that -gap by sending a GET to a path on the site's own domain, which CloudFront then writes into the -access log like any other request. +The browser beacon records events that CloudFront page requests miss, including SPA route +changes and events raised by the site. ```typescript import { startBeacon } from "@kensio/rainlytics/beacon"; @@ -11,93 +9,63 @@ import { startBeacon } from "@kensio/rainlytics/beacon"; const beacon = startBeacon(); ``` -That is the whole of the setup. There is no script tag, no second host and no extra connection. The -module is bundled into the site's own JavaScript, and it goes out in a download the page was already -making. +Import it into your site's existing JavaScript bundle. Deploy [Beacon path](../beacon-path/) first +so `/_rainlytics` returns 204 at the CloudFront edge. -Deploy [the collection path](../beacon-path/) first. Without it the site's origin answers these -requests, which is the one thing the design exists to avoid. +## Report events -## What it collects +`startBeacon` watches `history.pushState`, `history.replaceState` and browser navigation. It reports +a `route` event when the visible path changes. -Every event carries the same three parameters. Two more travel where the event has them. +CloudFront already logged the document request for the first page. The beacon begins with later +route changes, avoiding a duplicate view. -| Parameter | Always | Holds | -| --------- | ------ | ------------------------------------------------------ | -| `v` | yes | The envelope version, so an old row still reads later. | -| `e` | yes | What happened, such as `route` or `lcp`. | -| `p` | yes | The page it happened on, as a path. | -| `n` | no | A number the event measured, such as a vital's value. | -| `m` | no | Text the event carries, such as what an error said. | - -A route change reports itself under the name `route`. Anything else is the site's own call: +Report a custom event with `report`: ```typescript -beacon.report({ event: "signup", page: location.pathname }); -beacon.report({ event: "inp", page: location.pathname, value: 180 }); -``` - -`n` and `m` are left out of the query string entirely where an event has neither, so a route change -is the length it always was. - -The request carries no cookies. `credentials: "omit"` is on every send, which keeps the site's own -cookies out of a header that would be paid for on each event and could reach a log. The beacon -generates no identifier of any kind. [Counting visitors](../visitors/) covers what a visitor is -here, and it is computed from the access log rather than from anything the browser sends. +beacon.report({ + event: "signup", + page: location.pathname, +}); -`event` and `page` are the site's own values, and the beacon sends whatever it is handed. Keep -personal data out of both. Whatever they hold is written into `cs_uri_query` in the access log and -kept for as long as the log objects are. +beacon.report({ + event: "purchase-value", + page: location.pathname, + value: 49.95, +}); +``` -**The page the beacon starts on is not reported.** Loading it was a request, CloudFront recorded it, -and reporting it again would count one view twice in two questions that are meant to agree. What the -beacon adds is every route change after that one. A route change to the page already showing is left -alone as well, which is what keeps a router putting a filter in the query string from reading as a -second view. +Each request contains an envelope version, event name and page. Events can also carry one number or +one text value. Keep personal data out of event names, pages and messages. These values remain in +the raw log until its lifecycle expires them. -[Beacon events](../beacon-events/) has the rollup that reads these rows back, including the cap that -bounds a flood of them. +The browser sends no cookies. Requests use `fetch` with `credentials: "omit"` and `keepalive: true`. +The beacon creates no browser identifier. -## Core Web Vitals +## Collect Core Web Vitals -Behind an import of its own, so a site reporting route changes does not pay for it: +Vitals use a separate entry point: ```typescript -import { startBeacon } from "@kensio/rainlytics/beacon"; import { reportVitals } from "@kensio/rainlytics/beacon/vitals"; -const beacon = startBeacon(); reportVitals(beacon); ``` -Four measurements, each sent once, each carrying its number in `n`. - -| Event | Is | Read from | Sent | -| ------ | ---------------------------- | -------------------------- | ------------------- | -| `ttfb` | Time To First Byte, ms | the `navigation` entry | straight away | -| `fcp` | First Contentful Paint, ms | a `paint` entry | when it happens | -| `lcp` | Largest Contentful Paint, ms | `largest-contentful-paint` | when the page hides | -| `cls` | Cumulative Layout Shift | `layout-shift` entries | when the page hides | - -TTFB and FCP are final as soon as they happen. LCP and CLS are not final until the page stops -painting and stops moving, so both are held until `visibilitychange` reports the document hidden and -sent then. That is the moment `keepalive` on the send exists for. Every ordinary way of leaving a -page hides the document first, including following a link and closing the tab, and a page that is -never hidden reports neither. +Rainlytics reports: -CLS is scored on the worst session window rather than the sum of every shift. A window runs no -longer than five seconds and ends after a second without a shift. A page that shifts a little every -few seconds all day would otherwise score as though it had shifted once, enormously. A shift the -reader caused is left out, which is the layout responding rather than the layout misbehaving. +| Event | Measurement | Unit | Sent | +| ------ | ------------------------ | ------------ | ----------------------- | +| `ttfb` | Time to First Byte | milliseconds | when available | +| `fcp` | First Contentful Paint | milliseconds | when available | +| `lcp` | Largest Contentful Paint | milliseconds | when the document hides | +| `cls` | Cumulative Layout Shift | score | when the document hides | -Each observer asks for `buffered` entries, so a paint that happened before the site's bundle ran -still reports. Without that, every fast page would report nothing. +CLS uses the worst session window and ignores shifts after recent user input. LCP and CLS wait until +the document hides because they can change while the page remains visible. -**INP is not collected, and that is deliberate.** It is a Core Web Vital, and computing it means -grouping event-timing entries by `interactionId` and taking a high percentile of the result. A -version of that with a subtle mistake in it reports a plausible number rather than an obvious -failure, which is the failure this project is least able to detect. A site that wants INP runs the -`web-vitals` library itself and hands the number over: +INP calculation needs interaction grouping and percentile logic. Rainlytics leaves that calculation +to `web-vitals`, and a site can report the result: ```typescript import { onINP } from "web-vitals"; @@ -107,12 +75,11 @@ onINP(({ value }) => { }); ``` -That is also the measured trade. `web-vitals` covering LCP, CLS and INP bundles to 3209 bytes -gzipped. The four above cost 550. +The shipped `web-vitals` rollup currently reports TTFB, FCP, LCP and CLS. -## JavaScript errors +## Collect JavaScript errors -Behind an import of its own as well, and that is a privacy decision as much as a page weight one: +Error reporting also uses a separate entry point: ```typescript import { reportErrors } from "@kensio/rainlytics/beacon/errors"; @@ -120,36 +87,11 @@ import { reportErrors } from "@kensio/rainlytics/beacon/errors"; reportErrors(beacon); ``` -An uncaught exception reports as `error` and an unhandled promise rejection as `rejection`, each -carrying what it said in `m`. The page is read when the error happens, so an error in a single-page -app is reported against the route it happened on. Neither listener handles the error. The browser -still logs to the console and any other handler on the page still runs. - -[`javascript-errors`](../javascript-errors/) counts those rows by page and message, most reported -first. A deployment opts into that rollup when it opts into error reporting. - -**No stack.** A stack names the URL of every frame and often a good deal more, none of it fits in a -query string worth storing, and the name and message are what a rollup counting errors would group -by. The message is cut at 200 characters, because the whole query string is stored for as long as -the log objects are. +Uncaught exceptions use event name `error`. Unhandled promise rejections use `rejection`. The +message is limited to 200 characters and no stack trace is sent. -## What a site holding no personal data gets - -[Counting visitors](../visitors/) has the field set that delivers no viewer address. A deployment -running it holds no personal data, and the beacon can hand some back. This is what each part does -about that. - -- **Route changes and vitals cannot.** A path the site publishes and a number are not personal data, - whoever is reading. -- **`event` and `page` are the site's own values.** The beacon sends what it is handed. A router - that puts an account name in a path puts it in the log. -- **An error message is the risk.** It is the site's own text, written by the site's own code, and - nobody audits it for what it interpolates. This is why errors are behind an import rather than a - flag. Importing them is the decision. - -`redact` is where a site takes it back out. It runs on the whole message, before the 200-character -cut and not after it, so a pattern written for a whole address still matches one that would have -been cut in half. Answering `undefined` reports nothing for that error: +Error messages can contain email addresses, account names and other personal data. Redact them +before they reach the immutable log: ```typescript reportErrors(beacon, { @@ -157,94 +99,50 @@ reportErrors(beacon, { }); ``` -Nothing already written comes back out. The raw store is immutable and keeps whatever was written -into it until the [log bucket](../log-bucket/) expiry reaches it, so this is a decision to take -before turning error reporting on rather than after. - -## What it costs a page +Return `undefined` from `redact` to drop a message. Query-time cleanup is too late to remove the raw +value. -Each import is measured on a bundle of what a site actually writes, minified and gzipped. +## Consent and stopping -| What a site imports | Minified | Gzipped | Budget | -| ------------------- | -------- | ------- | ------ | -| The beacon | 1042 | 586 | 640 | -| With vitals | 2290 | 1136 | 1250 | -| With errors | 1489 | 756 | 880 | -| All of it | 2853 | 1349 | 1500 | - -`pnpm check` fails over any of those budgets. Brotli, which CloudFront serves to anything that asks, -comes in under the gzip figure. - -Vitals and errors are separate imports so that these are separate numbers. A site reporting route -changes alone pays the first row and nothing else. - -Bytes are not the only cost. The beacon adds no DNS lookup, no TLS handshake and no connection, -because the collection path is on the origin the page is already talking to. One event is a request -of a few hundred bytes that is answered at the edge with a 204 and no body. - -## The browser floor - -`browserslist` in package.json says `baseline 2022`, which is Chrome 108, Edge 108, Firefox 108 and -Safari 16. The build targets ES2021, and `tsconfig.json` has why that is a version behind. - -The floor is set low on purpose. A beacon that fails on an older browser annoys nobody, and that is -the problem with it. The visitor still reads the site and still appears in the CloudFront rows layer -1 counts, so the two layers disagree by an amount invisible from inside the data. A low floor costs -nothing here, since none of this code needs syntax newer than ES2015. - -Sending uses `fetch` with `keepalive`, which lets a request outlive the page that started it. Chrome -has had it since 66 and Safari since 13, and Firefox only since 133 (December 2024). A browser -without it ignores the option and sends the request anyway. What is lost there is the last event -before somebody navigates away, and a route change happens with the page still open. - -## Switching it off - -Consent belongs to the site. Nothing here reads a banner, a cookie or `navigator.doNotTrack`. - -Call `startBeacon` once somebody has agreed: +The site owns the consent decision. Start the beacon after the site's consent flow approves it: ```typescript const beacon = consented ? startBeacon() : undefined; ``` -And stop it if they take it back: +Stop reporting when consent is withdrawn: ```typescript beacon?.stop(); ``` -`stop` puts back the `History` methods that starting it wrapped, and `report` sends nothing -afterwards. Calling it twice is safe. - -A consent story built into the beacon would be one more thing every page downloads, and it would be -wrong for whichever banner the site actually runs. +`stop` restores the wrapped history methods and makes later `report` calls inert. ## Options -| Option | Default | What it changes | -| -------------- | -------------- | ------------------------------------------- | -| `path` | `/_rainlytics` | The collection path, matching `BeaconPath`. | -| `reportRoutes` | `true` | Whether route changes report themselves. | - -A site that passed `path` to the construct passes the same one here. If the two disagree, the beacon -reports to a path nothing answers, and the first sign of that is a dataset holding no beacon rows. - -`reportRoutes: false` leaves `report` as the only way an event is sent, which suits a site that would -rather call its own router's hook. +```typescript +const beacon = startBeacon({ + path: "/_measure", + reportRoutes: false, +}); +``` -## What it deliberately does not do +`path` defaults to `/_rainlytics` and must match `BeaconPath`. Set `reportRoutes: false` when your +framework already provides a router hook and you only want explicit `report` calls. -**No sampling.** A beacon event is a row in a log object the site is already paying for, and -`beacon-events` bounds a flood in the query. Sampling would cost bytes on every page to save nothing -worth saving, and it would put a scaling factor in front of numbers that are otherwise counts. +## Page weight and browser support -**No INP.** The section on vitals above has why, and what to do about it. +The base beacon is currently 586 bytes gzipped. Vitals and errors together bring the complete set +to 1,349 bytes gzipped. Project checks bundle, minify and gzip each entry point and fail when a size +budget is exceeded. -**No `navigator.sendBeacon`.** It is POST-only, and the whole design rests on a GET whose query -string CloudFront writes into `cs-uri-query`. +The browser target is Baseline 2022 (Chrome and Edge 108, Firefox 108 and Safari 16). A browser +without fetch keepalive still attempts the request, but its final event is more likely to be lost +during navigation. diff --git a/docs/command-line/README.md b/docs/command-line/README.md index 7e3da7b..87de92e 100644 --- a/docs/command-line/README.md +++ b/docs/command-line/README.md @@ -1,286 +1,153 @@ # Command line -`rainlytics` is how the data collected in your own AWS account is read back. There is no dashboard, -and this page explains why. +The `rainlytics` command reads analytics from your AWS account. ```bash -npx @kensio/rainlytics --help -``` - -That runs on a machine with only Node on it. `aws-cdk-lib` and `constructs` are optional peer -dependencies of this package, and an install that only wants the command line skips both. - -`--help` on the root command and on every subcommand is the documentation. A reader should find -everything there, and this page repeats it for people reading the site. - -The named questions read one rollup at a time. [`report`](#reading-a-calendar-report) reads a -versioned document containing several sections for one closed calendar period. -[`saved-query`](#running-a-query-saved-in-the-workgroup) runs a question a site saved for itself, -and [`query`](../query/) takes SQL for everything else. - -## Why a command line - -Authentication, mostly. A dashboard needs an identity story, a session model and a way to revoke -access, and that subsystem then has to stay secure for as long as it exists. A command line on the -AWS SDK's default credential chain inherits SSO, MFA, role assumption, least-privilege IAM policies -and CloudTrail audit of who queried what, and none of that is code in this repository. For a product -whose premise is that it runs in your own AWS account, your existing AWS setup is the coherent -answer. - -Structured output is the second reason. A person at a terminal and an assistant with shell access -run the same commands and read the same structured answer, with no API to build in between. - -## Authentication - -There is no Rainlytics account, password or API key. Credentials come from the AWS SDK's default -chain, which is the same one the AWS CLI reads: - -- An SSO session from `aws sso login`. -- A named profile in `AWS_PROFILE`, or the default profile in `~/.aws/credentials`. -- `AWS_ACCESS_KEY_ID` and friends in the environment. -- A role assumed through `AWS_ROLE_ARN`, including the one a CI runner is given. -- The credentials of the EC2 instance, the container or the Lambda it is running in. - -Whatever that chain resolves to is what the queries run as, and CloudTrail records them under that -identity. - -## Permissions - -The identity that chain resolves to needs the permissions for what the command does, and the two -halves of the command surface need different ones. - -A named question or calendar report reads a precomputed object. That takes `s3:GetObject` on the -summaries bucket, which an SSO read-only role already carries. - -`query`, `saved-query` and `--query` run Athena, which takes four more. Those are -`athena:StartQueryExecution` and `athena:StopQueryExecution` on the workgroup, and `s3:PutObject` -and `s3:AbortMultipartUpload` on the bucket that workgroup writes results to. A read-only role has -none of the four, and a command refused for want of them says so: - -```text -rainlytics: User: arn:aws:sts::000000000000:assumed-role/AWSReservedSSO_ReadOnly/... is not -authorized to perform: athena:StartQueryExecution on resource: -arn:aws:athena:eu-west-1:000000000000:workgroup/rainlytics -Running a query takes athena:StartQueryExecution and athena:StopQueryExecution on the rainlytics -workgroup, and s3:PutObject and s3:AbortMultipartUpload on the bucket that workgroup writes results -to. A named question (pageviews, referrers, browsers, status-codes, cache-hit-ratio or searches) -answers from a precomputed summary on s3:GetObject alone. Name the bucket holding those with ---summaries, or put it in RAINLYTICS_SUMMARY_BUCKET. +pnpm exec rainlytics --help ``` -The [query workgroup](../query-workgroup/) page has the whole policy, including the reads a -read-only role already allows. - -## Region - -Every command that reaches AWS takes `--region`: +You can also run the package without installing it in a project: ```bash -rainlytics pageviews --last 30d --region us-east-1 +npx @kensio/rainlytics --help ``` -Left off, the region comes from the same chain as the credentials. It reads `AWS_REGION`, then the -region set on the profile, then the instance metadata on EC2. Athena commands also take `--database` -and `--workgroup`. +Rainlytics has no account, password or API key of its own. -The region decides more than a default suggests. A workgroup, a Glue table and an S3 bucket each -exist in one region, and a query asked in another is answered `WorkGroup rainlytics is not found.` -about a workgroup sitting where it was deployed. Rainlytics adds the region it asked in to that -message. The [query](../query/) page has the whole failure, and why the region to ask is the one the -log bucket is in. +## Credentials and region -## Where an answer comes from +The command uses the AWS SDK default credential chain. It supports named AWS profiles, IAM Identity +Center sessions, assumed roles, environment credentials and workload roles. -The named questions read precomputed summaries off S3. A schedule counted each window once, the -command fetches the windows the range covers, and the whole read costs a GET each. `report` reads one -precomputed report document from the same bucket. `query` and `saved-query` run Athena because -ad-hoc SQL has no stored answer. +Choose a profile and region with the normal AWS settings: ```bash -rainlytics pageviews --last 7d --summaries rainlytics-summaries-1a2b +AWS_PROFILE=analytics AWS_REGION=us-east-1 rainlytics pageviews --last 7d ``` -`--summaries` names the bucket the [summary schedule](../summary-schedule/) writes to. -`RAINLYTICS_SUMMARY_BUCKET` in the environment says it once for a whole shell, and it is the same -variable `RollupSummaries` sets on its own job. With neither, the command says where to put it and -stops. +Every command that reaches AWS also accepts `--region`. Athena commands accept `--database` and +`--workgroup` when the deployment changed their defaults. -The [summary schedule](../summary-schedule/#give-the-command-line-the-generated-bucket-name) page -shows how to publish a generated bucket name through a CloudFormation output and have `cdk deploy` -write it to a local JSON file. Shell setup can read `RAINLYTICS_SUMMARY_BUCKET` from that file. The -command still needs only `s3:GetObject` and avoids a CloudFormation lookup on every run. +The region must contain the Glue table, Athena workgroup and summary bucket. A missing workgroup in +the selected region is often a profile or region mistake. -`--query` runs the question through Athena: +## Read a named question + +The named commands cover pageviews, referrers, browsers, status codes, cache hit ratio, searches, +JavaScript errors and Web Vitals. ```bash -rainlytics pageviews --last 7d --query +rainlytics pageviews --last 7d +rainlytics referrers --last 30d +rainlytics status-codes --last 24h --include-bots ``` -That answer is fresher than the last scheduled run, and it covers windows no schedule has computed. -Athena charges per byte scanned and the command says what it came to. Nothing reaches for it on its -own. A command that queried whenever a summary was missing would put that charge back without -anybody choosing it. +These commands read precomputed summaries. Name the bucket with an option or environment variable: -The rows are the same either way, and a pipeline reading the JSON sees no difference. What changes is -on standard error: +```bash +rainlytics pageviews --last 7d --summaries rainlytics-summaries-1a2b -```text -Read 23 summaries of pageviews from rainlytics-summaries-1a2b, covering -2026-08-21T15:00:00.000Z to 2026-08-28T14:00:00.000Z. -The newest was computed 2026-08-28T14:15:03.001Z (13 minutes ago). 23 GETs, -about $0.0000092 at the us-east-1 rate. +export RAINLYTICS_SUMMARY_BUCKET=rainlytics-summaries-1a2b +rainlytics pageviews --last 7d ``` -The span there is the one that answered, and it runs a little short of the one asked for. The hour -running now is still filling and has no stored window. Where the range reaches further back than the -schedule does, a further line says how many windows had no summary. - -[Rollups](../rollups/#reading-a-precomputed-answer) has which filters a stored summary can answer -under, and [rollup summaries](../summaries/#reading-one-back) has what happens to a range no stored -window covers. - -## Output - -`--output` takes `json`, `csv` or `table` for commands that answer with rows. - -Left off, it is `table` when standard output is a terminal and `json` when standard output is piped -or redirected. That is what makes a pipeline work as typed: +Use `--query` to run the same question against raw logs with Athena: ```bash -rainlytics # a table, for reading -rainlytics | jq '.[0]' # JSON, with no flag passed -rainlytics --output csv > views.csv # or ask for something else +rainlytics pageviews --last 2h --query ``` -**JSON** is an array of objects, one per row, with no envelope around it. Every object carries every -column, and a value the row left empty is `null`. So `.[0].path` is the expression, and a key stays -put all the way down the array. +This produces a fresh result and incurs Athena query cost. Rainlytics never falls back to Athena +automatically when a stored summary is missing. -**CSV** follows RFC 4180, with a header line even for a result of zero rows. A field carrying a -comma, a quotation mark or a newline is quoted, and a quotation mark inside one is doubled. Lines -end in LF where the RFC asks for CRLF, because everything that reads a CSV takes LF and the tools in -between mind a stray CR. +## Read a calendar report -**Table** pads each column to its widest value, rules the heading off, and adds no colour and no -trailing whitespace. +`report` reads one stored report for a closed calendar period: -`report` has JSON output only because its period, schema version, source coverage and section -metadata are part of the result. It writes the same JSON document at a terminal and in a pipe. -Passing `--output csv` or `--output table` is a usage error. +```bash +rainlytics report day 2026-08-30 +rainlytics report week 2026-08-24 --time-zone Europe/London +rainlytics report month 2026-07 --compare +rainlytics report year 2025 +``` -### CSV in a spreadsheet +The time zone and first weekday must match the `RollupSummaries` deployment. Defaults are UTC and +Monday. `--compare` reads the preceding period too and calculates the changes between both stored +reports. Report commands never run Athena. -Referrers and user agents are written by whoever made the request. Excel and LibreOffice treat a -field opening with `=`, `+`, `-` or `@` as a formula, so a crafted referrer can become one when the -file is double-clicked. +Reports always use JSON because their period, source coverage and section metadata are part of the +answer. -Rainlytics writes the value exactly as it read it, because the same file is what a script -downstream parses. Where the data is going anywhere near a spreadsheet, import the file rather than -double-clicking it, and choose text for those columns. +## Run saved SQL -## Streams and exit codes +`RollupQueries` stores generated SQL in Athena. Run a saved query by name: -The result of a command goes to standard output. Everything else, including help, warnings and the -reason for a failure, goes to standard error. A pipeline therefore reads data and never prose. +```bash +rainlytics saved-query countries +``` -| Exit | Meaning | -| ---- | ------------------------------------------------------------------- | -| 0 | It worked. | -| 1 | The command ran and could not finish. A retry sometimes gets past. | -| 2 | The command line was wrong. Running it again unchanged fails again. | +The `rainlytics-` prefix is optional. `countries` and `rainlytics-countries` select the same saved +query. -Two failure codes because they call for different responses, and 2 for a usage error is the -convention `getopt` set and Python's `argparse` kept. +A saved query already contains its range, row limit and filters. It accepts `--region`, +`--workgroup` and output options, but it does not accept named-rollup filters such as `--last` or +`--path`. -Standard output stays empty on both. Nothing that failed writes a partial result that a later step -could mistake for a whole one. +Use `rainlytics query` for SQL written at the terminal. See [Query](../query/). -## Reading a calendar report +## Output formats -`rainlytics report` reads one closed day, week, month or year from the summaries bucket: +Named questions, saved queries and ad-hoc queries support JSON, CSV and tables. ```bash -rainlytics report day 2026-08-30 -rainlytics report week 2026-08-24 --time-zone Europe/London -rainlytics report month 2026-07 --compare -rainlytics report year 2025 +rainlytics pageviews --last 7d --output table +rainlytics pageviews --last 7d --output json +rainlytics pageviews --last 7d --output csv > pageviews.csv ``` -The date selects the period. A weekly date can be any date inside the week. `--time-zone` defaults -to UTC and must match the value passed to `RollupSummaries`. For a weekly report, -`--week-starts-on` defaults to Monday and must also match. These values are part of the report's S3 -address, and the command derives the key from them. - -`--summaries` names the bucket, and `RAINLYTICS_SUMMARY_BUCKET` supplies the same default used by -the named questions. `--region` uses the AWS SDK credential and region chain when left off. - -The versioned report document is the whole of standard output. Standard error names the bucket and -object key, the object's last-modified age and the price of one S3 GET. A missing, incomplete or -unsupported document exits non-zero with empty standard output. Reading a report never starts an -Athena query. - -`--compare` derives changes against the immediately preceding calendar period. It reads the earlier -stored report with one additional S3 GET and writes a versioned comparison document. Both report -periods, computation times and source coverage values stay in the result. Standard error names both -object keys and the price of two GETs. The [calendar reports](../reports/#comparing-adjacent-periods) -page defines the metric and missing-data rules. - -## Running a query saved in the workgroup - -`rainlytics saved-query` runs a query Athena already holds, by name: +Without `--output`, a terminal receives a table and a pipe receives JSON: ```bash -rainlytics saved-query countries +rainlytics pageviews --last 7d | jq '.[0].path' ``` -The [`RollupQueries`](../rollups/#the-same-sql-saved-in-the-console) construct saves one named query -per rollup, and a site that [wrote a rollup of its own](../rollups/#writing-a-rollup-of-your-own) -saves that beside them. That is how a question this package never shipped gets a command line, with -the `--output` formats, the cost report and the exit codes every other command has. +JSON output is an array of row objects. CSV follows RFC 4180 apart from using LF line endings. Table +output contains no color codes or trailing spaces. + +Import CSV into spreadsheet software with every external-text column set to text. Referrers and +user agents come from requests, and a value beginning with `=`, `+`, `-` or `@` can be treated as a +formula when a CSV file is opened directly. -The name is the one Athena lists, with or without the `rainlytics-` prefix the construct adds. -`countries` and `rainlytics-countries` reach the same saved query. A name matching nothing is -answered with the names that are saved in the workgroup. A guess is one way to find out what is -there. +## Output streams and exit codes -What a saved query covers was settled when it was saved. `--last`, `--limit`, `--include-bots`, -`--path`, `--host` and `--param` are absent here, and each is refused rather than accepted and -ignored. The SQL Athena holds carries a range and a row count already, and a saved rollup covers the -month you run it in. `requests` on the construct is where the rest of it is decided. +Data goes to standard output. Help, warnings, source coverage and query cost go to standard error. +Standard output stays empty when a command fails. -The database comes from the saved query too. Athena records the one a query was written against, and -this runs it against that one. `--workgroup` and `--region` say where to look, and the query runs -where it was found. +| Exit code | Meaning | +| --------- | --------------------------------------------------------------- | +| `0` | The command completed. | +| `1` | AWS or the requested operation failed. A retry may succeed. | +| `2` | The command line was invalid. The same command will fail again. | -## A command comes before its options +Put the command before its options: ```bash -rainlytics --output csv # this way round -rainlytics --output csv # refused, with that sentence +rainlytics pageviews --output csv ``` -`rainlytics --help` and `rainlytics --version` are the only lines with no command in them. +Run `rainlytics --help` for the full option list and examples for one command. -## What it can do today - -Answer the six default questions with [`rainlytics pageviews`, `referrers`, `browsers`, -`status-codes`, `cache-hit-ratio` and `searches`](../rollups/). A deployment using the browser -beacon can add the shipped [`javascript-errors`](../javascript-errors/) and -[`web-vitals`](../web-vitals/) questions. Read a closed day, week, month or year with -[`rainlytics report`](#reading-a-calendar-report). -Run a question a site saved for itself with -[`rainlytics saved-query`](#running-a-query-saved-in-the-workgroup), and run SQL for anything else -with [`rainlytics query`](../query/). +## Permissions -The named questions read [precomputed summaries](../summaries/) off S3. A calendar report is one -precomputed document. `query` and `saved-query` reach Athena, and so does a named question given -`--query`. +Reading named questions and reports needs `s3:GetObject` on the summaries bucket. A normal AWS +read-only role often has this access. -Rainlytics is experimental and pre-1.0. The command surface will change without a major version -behind it. +Athena operations need query access to the workgroup, Glue catalog reads, raw log reads and writes +to the query results bucket. Grant the set with `workgroup.grantQuerying(role, table)`. See [Query +workgroup](../query-workgroup/#grant-query-access). diff --git a/docs/getting-started/README.md b/docs/getting-started/README.md new file mode 100644 index 0000000..47e6977 --- /dev/null +++ b/docs/getting-started/README.md @@ -0,0 +1,217 @@ +# Getting started + +This guide deploys Rainlytics for one existing CloudFront distribution and reads the first stored +pageview report. + +You need: + +- Node.js 22 or newer +- an AWS CDK app written in TypeScript +- a deployed CloudFront distribution +- AWS credentials that can deploy the resources in this guide + +Rainlytics configures CloudFront standard logging v2 through the CloudWatch Logs API in +`us-east-1`. The example keeps the log bucket, Glue table, Athena workgroup and scheduled jobs in +that region too. + +## Install Rainlytics + +Add Rainlytics to your CDK app. Most CDK apps already have the two peer dependencies. + +```bash +pnpm add @kensio/rainlytics aws-cdk-lib constructs +``` + +The package also installs the `rainlytics` command. + +## Create the visitor salt + +The default pageview rollup counts visitors. It derives daily identifiers from a secret stored as +an SSM Parameter Store `SecureString`. + +Create the secret once in the account and region where the scheduled jobs will run: + +```bash +aws ssm put-parameter \ + --region us-east-1 \ + --name /rainlytics/visitor-salt \ + --type SecureString \ + --value "$(openssl rand -hex 32)" +``` + +Use your normal AWS CLI profile or role for this command. Rainlytics never writes the secret to a +CloudFormation template. Keep the parameter after deployment because recomputing an old period +requires the same secret. + +You can omit this step by excluding viewer addresses from log delivery. See [Counting +visitors](../visitors/#run-without-visitor-counts). + +## Add an analytics stack + +Create a stack like this in your CDK app. Replace the distribution ID with your own. + +```typescript +import { App, CfnOutput, Stack, type StackProps } from "aws-cdk-lib"; +import { Construct } from "constructs"; + +import { + CloudFrontLogDelivery, + LogBucket, + LogTable, + QueryWorkgroup, + RollupQueries, + RollupSummaries, +} from "@kensio/rainlytics/cdk"; + +class AnalyticsStack extends Stack { + constructor(scope: Construct, id: string, props: StackProps) { + super(scope, id, props); + + const logs = new LogBucket(this, "Logs"); + + const delivery = new CloudFrontLogDelivery(this, "Delivery", { + distributionId: "E1EXAMPLE1234", + logBucket: logs.bucket, + }); + + const table = new LogTable(this, "Table", { + deliveries: [delivery], + }); + + const workgroup = new QueryWorkgroup(this, "Workgroup"); + + new RollupQueries(this, "SavedQueries", { table, workgroup }); + + const summaries = new RollupSummaries(this, "Summaries", { + table, + workgroup, + }); + + new CfnOutput(this, "SummaryBucketName", { + value: summaries.bucket.bucketName, + }); + } +} + +const app = new App(); + +new AnalyticsStack(app, "Analytics", { + env: { + account: process.env["CDK_DEFAULT_ACCOUNT"], + region: "us-east-1", + }, +}); +``` + +This stack creates: + +- a private, versioned S3 bucket for raw logs +- a CloudFront log delivery with hourly Hive partitions +- a Glue database and projected table +- an Athena workgroup with a per-query scan limit +- saved versions of the built-in rollup queries +- scheduled summary and calendar-report jobs +- a second S3 bucket for stored answers + +The log bucket is the source of record. Its objects are retained for 370 days by default. The +summary and query result buckets are separate because they hold derived data with different +retention rules. + +## Synthesize and deploy + +Check the template before deployment: + +```bash +pnpm exec cdk synth Analytics +pnpm exec cdk diff --method=template Analytics +pnpm exec cdk deploy Analytics +``` + +The deploy prints `SummaryBucketName`. Keep that value for the command line. + +CloudFront can take up to 12 hours to apply a logging change. New log objects then appear under a +path like this: + +```text +s3:///rainlytics/distributionid=E1EXAMPLE1234/year=2026/month=09/day=01/hour=14/ +``` + +The first scheduled summary is written after a complete hour closes and CloudFront delivers its +logs. A new deployment therefore has no immediate historical summaries. + +## Run the command line + +Set the region and summary bucket in your shell: + +```bash +export AWS_REGION=us-east-1 +export RAINLYTICS_SUMMARY_BUCKET= +``` + +Run a named question: + +```bash +pnpm exec rainlytics pageviews --last 24h +``` + +The command reads your AWS credentials from the standard SDK credential chain. It prints a table at +a terminal and JSON when piped. + +```bash +pnpm exec rainlytics pageviews --last 24h | jq '.[0]' +pnpm exec rainlytics status-codes --last 24h --output csv > status-codes.csv +``` + +Named questions read the precomputed objects in the summary bucket. Add `--query` to calculate the +answer directly from the raw logs with Athena: + +```bash +pnpm exec rainlytics pageviews --last 2h --query +``` + +An Athena query needs write access to the workgroup's results bucket and query permissions that a +read-only role usually lacks. The [Query workgroup](../query-workgroup/#grant-query-access) page +shows how to grant the complete set from CDK. + +## Add the browser beacon + +The access-log pipeline is complete at this point. Add the browser module only when you need SPA +route changes, custom events, Web Vitals or JavaScript errors. + +`BeaconPath` must be added where your CDK app has the `Distribution` and its origin: + +```typescript +import { BeaconPath } from "@kensio/rainlytics/cdk"; + +new BeaconPath(this, "AnalyticsBeacon", { + distribution, + origin, +}); +``` + +Start the browser module in your site's existing JavaScript bundle: + +```typescript +import { startBeacon } from "@kensio/rainlytics/beacon"; + +const beacon = startBeacon(); +beacon.report({ event: "signup", page: location.pathname }); +``` + +The collection path defaults to `/_rainlytics`. The CDK construct and browser module must use the +same path. Continue with [Browser beacon](../beacon/) for Core Web Vitals, error reporting, consent +and custom event rollups. + +## Next steps + +- Read [Rollups](../rollups/) for every built-in question and filter. +- Read [Command line](../command-line/) for profiles, output formats and reports. +- Adjust raw log retention in [Log bucket](../log-bucket/). +- Review visitor data handling in [Counting visitors](../visitors/). +- Review costs and failure checks in [Summary schedule](../summary-schedule/). + + diff --git a/docs/javascript-errors/README.md b/docs/javascript-errors/README.md index b62a3d0..6211f46 100644 --- a/docs/javascript-errors/README.md +++ b/docs/javascript-errors/README.md @@ -1,13 +1,12 @@ # JavaScript errors -`javascript-errors` counts uncaught exceptions and unhandled promise rejections by page and -message, most reported first. +The `javascript-errors` command counts uncaught exceptions and unhandled promise rejections by page +and message. ```typescript import { javascriptErrors, rollups } from "@kensio/rainlytics"; -import { RollupSummaries } from "@kensio/rainlytics/cdk"; -new RollupSummaries(this, "RainlyticsSummaries", { +new RollupSummaries(this, "Summaries", { table, workgroup, rollups: [...rollups, javascriptErrors], @@ -23,43 +22,32 @@ page message errors --------- ------------------------------------------------ ------ /checkout TypeError: Cannot read properties of undefined 27 /account Error: Session expired 9 -/checkout Error: Order 41 failed 1 ``` -The count covers the `error` rows produced by uncaught exceptions and the `rejection` rows produced -by unhandled promise rejections. Route changes, Web Vitals and custom events are left out. +## Enable collection -## A site opts into it +Import error reporting in the measured site: -The six default rollups work from access-log fields every deployment has. This one reads rows from -optional browser error reporting. A site without those rows would pay for an empty Athena query on -every window, so `javascriptErrors` stays outside the exported `rollups` list. +```typescript +import { startBeacon } from "@kensio/rainlytics/beacon"; +import { reportErrors } from "@kensio/rainlytics/beacon/errors"; -Adding it on both cadences with the default two-window recomputation makes 50 more Athena queries a -day. At Athena's ten million byte minimum, that comes to about 8 cents a month. Lambda and S3 add a -few cents or less at this scale. +const beacon = startBeacon(); +reportErrors(beacon); +``` -The command ships for every deployment. A missing summary is reported as a window that was never -computed. `--query` runs the same question from raw logs. +`error` events come from uncaught exceptions. `rejection` events come from unhandled promise +rejections. The browser's normal error behavior is unchanged. -## Messages stay exact +The rollup is optional because a site without browser error events would pay for an empty query on +every scheduled window. Adding it under the default schedule adds 50 Athena queries a day. -The message is grouped exactly as `reportErrors` sent it, after the site's `redact` function and the -200-character limit. The group key contains the page and message only. An exception and a rejection -with the same values therefore share one row. +## Grouping and redaction -An interpolated value produces a group for each value. These messages are separate: +Rows group by the exact page and message sent by the browser. Error and rejection events with the +same page and message share one row. Interpolated identifiers create separate groups. -```text -Error: Order 41 failed -Error: Order 42 failed -``` - -Rainlytics keeps every part of the message. A number can be an order identifier, an HTTP status or a -line number that separates two failures. A general replacement would merge some errors that need -different fixes. - -A site that wants one group can normalise the message before sending it: +Normalize a message before sending when several values should form one group: ```typescript reportErrors(beacon, { @@ -67,51 +55,35 @@ reportErrors(beacon, { }); ``` -That replacement also changes the immutable raw row. Use `rainlytics query` or a custom rollup with -`regexp_replace` when the stored message must retain the value. +This also protects the raw log. A query-time replacement changes the report but leaves the original +message in S3. -## The page comes from the event +Error messages can contain personal data. Review every message your application can produce and use +`redact` before enabling collection. Messages are limited to 200 characters and stack traces are +never sent. -The request itself goes to the collection path. Its `p` parameter records the page that was visible -when the error happened, including a route reached without a document request in a single-page app. -The rollup groups on that decoded value. +## Custom collection paths -The collection path defaults to `/_rainlytics`, matching `BeaconPath`. A deployment using another -path records it on the summaries: +The rollup reads the default `/_rainlytics` path. Record another path in its summary request: ```typescript -new RollupSummaries(this, "RainlyticsSummaries", { +new RollupSummaries(this, "Summaries", { table, workgroup, rollups: [...rollups, javascriptErrors], - requests: { "javascript-errors": { paths: ["/_measure"] } }, + requests: { + "javascript-errors": { paths: ["/_measure"] }, + }, }); ``` -The matching fresh query is: - -```bash -rainlytics javascript-errors --path /_measure --last 7d --query -``` - -## Messages can contain personal data - -An error message is text written by the measured site's own code. It can contain an email address, -an account name or another value about the person using the page. The rollup copies that message -into its summary and the command prints it. - -Read the [browser beacon](../beacon/#what-a-site-holding-no-personal-data-gets) page before enabling -error reporting. Its `redact` option runs before the message reaches the access log. Query-time -normalisation leaves the value in the immutable raw row. Privacy filtering belongs in `redact`. - -## Combining stored windows +Use the same path with `BeaconPath` and `startBeacon`. -Error counts add across stored windows. Rows match on both page and message. The command orders the -combined counts again and applies its requested limit. +## Combined windows -Each summary stores only the leading rows from its own window. An error outside every stored top -list cannot appear in a combined answer, even when its total would be high enough. The command marks -that ranking as approximate. Use `--query` to rank all raw rows over the requested span. +Error counts add across stored windows. Rows match on page and message. Rankings across several +windows are approximate because each summary only stores its leading rows. Add `--query` to rank +all raw events across the full range. diff --git a/docs/log-delivery/README.md b/docs/log-delivery/README.md index 567ef01..d477fa3 100644 --- a/docs/log-delivery/README.md +++ b/docs/log-delivery/README.md @@ -1,275 +1,122 @@ # Log delivery -Sends a CloudFront distribution's access logs into a Rainlytics log bucket, partitioned and with the -field set the rollups read. This is the piece that turns a bucket into a pipeline. +`CloudFrontLogDelivery` sends standard logging v2 records from a CloudFront distribution to a +`LogBucket`. ```typescript import { CloudFrontLogDelivery, LogBucket } from "@kensio/rainlytics/cdk"; -const logs = new LogBucket(this, "RainlyticsLogs"); +const logs = new LogBucket(this, "Logs"); -new CloudFrontLogDelivery(this, "RainlyticsDelivery", { +const delivery = new CloudFrontLogDelivery(this, "Delivery", { distributionId: "E1EXAMPLE1234", logBucket: logs.bucket, }); ``` -## It has to live in us-east-1 +Pass `delivery` to `LogTable`. -Standard logging v2 is configured through the CloudWatch Logs API, and that API only accepts these -calls in us-east-1 whatever region the bucket is in. The construct refuses to synthesise anywhere -else, naming the stack. +## Deploy from `us-east-1` -Most sites keep their distribution somewhere else. That usually means a second stack: +AWS configures CloudFront standard logging v2 through the CloudWatch Logs API in `us-east-1`. +Rainlytics refuses to synthesize this construct in another region. + +The distribution can be defined in another stack. The delivery only needs its ID: ```typescript -const delivery = new Stack(app, "RainlyticsDeliveryStack", { - env: { account, region: "us-east-1" }, +new CloudFrontLogDelivery(this, "Delivery", { + distributionId: distribution.distributionId, + logBucket: logs.bucket, }); ``` -That is also why `distributionId` is a string rather than an `IDistribution`. Passing a construct -across regions needs CDK's cross-region references and the custom resources they bring with them. A -literal id needs neither. Pass `distribution.distributionId` if the two stacks are arranged so it -resolves. +Using a string avoids a CDK cross-region reference when the distribution belongs to another stack. + +A distribution can have one standard logging v2 delivery source. Remove an existing v2 source +before adding this construct. Legacy standard logging is separate and can continue at the same +time. -## One delivery source per distribution +CloudFront may take up to 12 hours to apply a logging change. -A distribution can carry a single delivery source. A second fails with `This ResourceId has already -been used in another Delivery Source in this account`. A distribution already running standard -logging v2 has to give that up before Rainlytics can take it over. Standard logging (legacy) is -separate and can keep running alongside. +## Output layout -## Logging changes take up to twelve hours +The defaults use JSON, hourly partitions and the `rainlytics` prefix: -A successful deploy is not the same as logs arriving. CloudFront applies a logging change within -twelve hours. An empty bucket the morning after a deploy is the normal case at that point. +```text +s3:///rainlytics/distributionid=E1EXAMPLE1234/year=2026/month=09/day=01/hour=14/ +``` -## Format, fields and partitions +The partition path is Hive-compatible. `LogTable` projects these keys and needs no crawler. -Output is JSON. Parquet is the better shape for a dataset Athena reads, and it carries a CloudWatch -conversion charge that AWS documents only by name. It stays an opt-in until there are numbers to -justify it. +Change the format or partition size when you create the delivery: ```typescript -new CloudFrontLogDelivery(this, "RainlyticsDelivery", { +const delivery = new CloudFrontLogDelivery(this, "Delivery", { distributionId: "E1EXAMPLE1234", logBucket: logs.bucket, outputFormat: "parquet", granularity: "daily", + prefix: "analytics", }); ``` -The output format can only be set when the delivery destination is created. Changing it later -replaces the destination rather than updating it. - -Fields default to the Rainlytics set, which is twelve fields and the minimum the rollups need. -Partitions are hourly by default, Hive-compatible, and land under a `rainlytics` prefix inside the -bucket. Changing the prefix later splits the dataset, because what was already written stays where -it was written. - -Objects arrive under the partition keys CloudFront derives from that layout: - -```text -s3://your-log-bucket/rainlytics/distributionid=E1EXAMPLE1234/year=2026/month=08/day=25/hour=14/ -``` - -Rainlytics sends the suffix path as bare variables (`{distributionid}/{yyyy}/{MM}/{dd}/{HH}`). -CloudFront supplies the `year=` half of each segment itself, because the delivery carries the -Hive-compatible option. A suffix path that has already added those key names is refused outright, -with `Provided suffixPath is invalid`. +JSON is the default because CloudFront charges for conversion to Parquet and the project does not +assume that conversion saves money. Parquet reduces Athena bytes scanned on larger datasets. Choose +the format from measured traffic and query cost. -## The field set holds the viewer's address +Changing the prefix or format later creates a second dataset shape. Existing objects stay at their +old keys and in their old format. -`c-ip` is one of the twelve. Rainlytics counts unique visitors as a hash of the viewer's address and -their user agent, under a salt that rotates every day, and a scheduled rollup computes that hash -from the address the log already holds. The reasoning is on -[#53](https://github.com/KensioSoftware/rainlytics/issues/53), in the comments, and -[Counting visitors](../visitors/) has what the number means and where the salt lives. +## Delivered fields -The raw store is therefore a record of people as well as of requests. Hashing downstream leaves the -addresses where they landed. CloudFront writes an object once and leaves it alone, and the addresses -last exactly as long as the log objects do. On the defaults that is 370 days, plus the 30 days a -superseded version survives, and the [log bucket](../log-bucket/) page has both numbers and how to -change them. +Rainlytics requests the smallest field set used by its built-in questions. The fields include the +request path, query string, referrer, user agent, host, country, response status, content type, +cache result, timestamp and viewer address. -A site that would rather not keep addresses delivers the named set without it: +The viewer address supports visitor counts and the repeated-event cap. To omit it: ```typescript import { logFieldNamesWithoutAddress } from "@kensio/rainlytics"; -new CloudFrontLogDelivery(this, "RainlyticsDelivery", { +const delivery = new CloudFrontLogDelivery(this, "Delivery", { distributionId: "E1EXAMPLE1234", logBucket: logs.bucket, fields: logFieldNamesWithoutAddress, }); ``` -That is the only line a site changes. The [log table](../log-table/) declares the columns the -delivery asks for and gets no `c_ip`. The [summary schedule](../summary-schedule/) reads the table, -computes the same six questions with the visitor count off, and is granted no permission to read -the salt. The SSM parameter a default deployment needs before its first run never has to exist. - -Pageviews, referrers, devices, status codes and geography all carry on. The visitor count is the one -thing that stops being computable, and no later job can recover it for the days the field was -absent. Geography survives because CloudFront resolves `c-country` at the edge from an address the -log then never records. - -## The first visitor counts cover part of a day - -A logging change takes up to twelve hours to apply, and it covers what CloudFront writes from then -on. The objects already in the bucket keep the shape they were written with, and the address is -absent from every record in them. A visitor count over a day that spans the change covers only the -part of the day with addresses in it. It reads low, and the answer looks like any other. Give it a -full day of delivered records before reading one day against another. - -## Encrypted buckets - -A bucket encrypted with a customer-managed key needs the delivery service allowed to use it. Without -that, the write is refused and the refusal appears in no log the bucket keeps. The construct adds -that grant, scoped to your account and its delivery sources, whenever the key is one CDK can reach. - -An imported key is the exception. CDK cannot add a statement to a key policy belonging to another -template. The grant would be written and never applied. The construct emits a build warning instead, -and the statement to add by hand is on the [log bucket](../log-bucket/) page. - -## Permissions for a scoped deploy role - -`cdk bootstrap` gives the CloudFormation execution role `AdministratorAccess` unless told otherwise. -An account still on that default deploys this construct with no IAM work at all, and can stop -reading here. The rest of this section is for a role narrowed with `cdk bootstrap ---cloudformation-execution-policies`, where a missing action arrives as a rolled-back deploy. - -Three `AWS::Logs::*` resources are created here (a delivery source naming the distribution, a -delivery destination naming the bucket and prefix, and a delivery joining the two), and one -CloudFront permission is checked on the caller. Establishing that list took three failed deploys on -the first site to run a narrowed role. The three headings below are the three failures, in the order -they arrived. - -### The CloudWatch Logs half - -```typescript -import { PolicyStatement } from "aws-cdk-lib/aws-iam"; - -new PolicyStatement({ - sid: "TheLogDelivery", - actions: [ - "logs:PutDeliverySource", - "logs:GetDeliverySource", - "logs:DeleteDeliverySource", - "logs:DescribeDeliverySources", - "logs:PutDeliveryDestination", - "logs:GetDeliveryDestination", - "logs:DeleteDeliveryDestination", - "logs:DescribeDeliveryDestinations", - "logs:PutDeliveryDestinationPolicy", - "logs:GetDeliveryDestinationPolicy", - "logs:DeleteDeliveryDestinationPolicy", - "logs:CreateDelivery", - "logs:GetDelivery", - "logs:DeleteDelivery", - "logs:DescribeDeliveries", - "logs:UpdateDeliveryConfiguration", - "logs:TagResource", - "logs:UntagResource", - "logs:ListTagsForResource", - ], - resources: ["*"], -}); -``` - -`"*"` is what has been deployed, and it is wider than it has to be. CloudWatch Logs supports -resource-level permissions on most of these actions, against `delivery-source`, -`delivery-destination` and `delivery` ARNs. A `Put` names the resource it is about, and a wildcard -such as `arn:aws:logs:us-east-1::delivery-source:*` therefore matches the call that creates -one. - -Narrowing it is more work than that makes it sound. `DescribeDeliverySources` and -`DescribeDeliveries` support no resource type and stay on `"*"`. `CreateDelivery` is authorised -against all three resource types at once, and so are the tagging actions. -Check each action in the [service authorization -reference](https://docs.aws.amazon.com/service-authorization/latest/reference/list_amazoncloudwatchlogs.html) -before scoping, and deploy the result before believing it. Nothing here has been run against a -scoped version of this statement. - -### Creates alone leave the stack stuck - -The reads and the deletes are the half worth being deliberate about, and the half a hand-written -policy drops. CloudFormation reads with `Get` before every call it makes. A policy carrying only the -`Put` actions therefore fails before it has created anything. - -Rollback then calls `Delete`, fails there too, and leaves the stack in `ROLLBACK_FAILED`. Clearing -that takes a hand `aws cloudformation delete-stack`. One missing verb family turns a failed deploy -into a stuck stack. - -### The distribution has to allow it - -Creating an `AWS::Logs::DeliverySource` calls CloudWatch Logs, and CloudWatch Logs then checks the -caller against CloudFront for the resource being logged. The denial therefore arrives from the -CloudWatch Logs API naming a CloudFront action: - -```text -User: .../cdk-hnb659fds-cfn-exec-role--us-east-1/AWSCloudFormation is not -authorized to perform: cloudfront:AllowVendedLogDeliveryForResource on resource: -arn:aws:cloudfront:::distribution/E1EXAMPLE1234 -(Service: CloudWatchLogs, Status Code: 400) -``` - -```typescript -new PolicyStatement({ - sid: "LogTheDistribution", - actions: ["cloudfront:AllowVendedLogDeliveryForResource"], - resources: [`arn:aws:cloudfront::${account}:distribution/E1EXAMPLE1234`], -}); -``` - -Scoped to the distribution. That is unusual for a `cloudfront:` action, and possible here because -the distribution has been serving the site for as long as it takes to know its id. A CloudFront ARN -carries no region. - -This one was reasoned about first and got wrong. The argument ran that v2 names the distribution by -ARN and CloudWatch Logs checks ownership on its own side, leaving the caller with no reason to hold -a `cloudfront:` action. A deploy said otherwise. - -The lesson generalises past this action. A permission that one service checks on another's behalf -cannot be worked out from the API surface, because the API being called is the wrong place to look -for it. Read the denial and grant what it names. - -### The bucket the delivery writes into needs its own permissions +`LogTable` follows the selected field set. `RollupSummaries` then disables visitor counts and does +not read the visitor salt. The other default questions continue to work. -The third deploy failed on `s3:CreateBucket`. That is the permission a reader of this page is most -likely to miss. Nothing above mentions S3, and a role assembled from this section alone gets as far -as creating the bucket and stops. +The first day after a field change contains a mixture of old and new records. Wait for a full day +of the new field set before comparing visitor totals. -[`LogBucket`](../log-bucket/) is a separate construct in the same stack, and the deploying role -needs the bucket verbs as well as the delivery ones. They are on the [log bucket](../log-bucket/) -page, along with a warning about scoping them to a generated bucket name. +## Customer-managed encryption -### Widening the policy can be a deploy of its own +When the log bucket uses a customer-managed KMS key, the delivery service needs permission to use +that key. Rainlytics adds the grant when the key belongs to the same CDK construct graph. An +imported key may need its owning stack updated. -Where the execution policy is itself managed by CDK, it usually lives in a different stack from the -one it governs, and sometimes in a different region. Editing it and rerunning the deploy that needed -it then changes nothing, because the policy stack has to go first. +Test delivery after deployment. An encryption-policy error prevents new objects from reaching S3, +and the destination bucket cannot log the failed write. -The failure that follows looks the same as the one just fixed. A reader who has added the missing -action and redeployed can reasonably conclude the action was wrong, when the policy carrying it was -never applied. +## Deployment permissions -### SSE-KMS is unverified +The standard CDK bootstrap role has enough permission. A restricted CloudFormation execution role +needs: -Passing `encryptionKey` puts a customer-managed key in the stack, and the construct grants -`delivery.logs.amazonaws.com` the use of it (see above). Whether the _deploying role_ needs `kms:` -actions of its own for that has not been established either way. The only consumer so far encrypts -with S3-managed keys and has no evidence about the path. +- CloudWatch Logs delivery source, destination and delivery management actions +- `cloudfront:AllowVendedLogDeliveryForResource` on the distribution +- the S3 permissions required by `LogBucket` +- IAM tagging and read/delete counterparts used by CloudFormation -The expectation, untested, is `kms:PutKeyPolicy` and `kms:GetKeyPolicy` on the key (CloudFormation -applies the grant by updating the key's policy), plus `kms:CreateKey`, `kms:DescribeKey`, -`kms:ScheduleKeyDeletion` and the tagging actions where the key is created in the same stack. Treat -that as a starting point for reading a denial, and not as a working policy. +Include `Get` and `Delete` actions as well as creation actions. CloudFormation reads resources +before updates and needs delete permission during rollback. A role with create-only permissions can +leave the stack in `ROLLBACK_FAILED`. diff --git a/docs/query/README.md b/docs/query/README.md index 88c2cf0..f84a71c 100644 --- a/docs/query/README.md +++ b/docs/query/README.md @@ -1,214 +1,99 @@ # Query -`rainlytics query` takes SQL, runs it through the Rainlytics workgroup, and prints the rows. +`rainlytics query` runs SQL through the Rainlytics Athena workgroup and prints every result row. ```bash rainlytics query "SELECT cs_uri_stem, count(*) AS views FROM cloudfront_logs - WHERE year = '2026' AND month = '08' AND day = '27' + WHERE year = '2026' AND month = '09' AND day = '01' GROUP BY 1 ORDER BY 2 DESC LIMIT 5" ``` -```text -cs_uri_stem views ---------------- ----- -/ 412 -/liju/ 208 -/grammar/ 97 -/pinyin/ 61 -/tones/ 44 -``` - -The SQL is one argument and has to be quoted. A shell splits an unquoted query on spaces and eats -the quotes inside it, which leaves Athena a different question from the one that was asked. The -command refuses a line that arrived in pieces. Running the first word of one would ask Athena -something else entirely and answer it. +Quote the SQL as one shell argument. -## What it needs +## Select the deployed resources -The [log table](../log-table/) and the [query workgroup](../query-workgroup/), both deployed. The -command reads their names from the same exported definition the constructs create them under. A -default deployment needs no flags: +The defaults use database `rainlytics`, table `cloudfront_logs` and workgroup `rainlytics`. ```bash -rainlytics query "SELECT count(*) FROM cloudfront_logs" +rainlytics query "SELECT count(*) FROM cloudfront_logs" \ + --region us-east-1 \ + --database rainlytics \ + --workgroup rainlytics ``` -Pass `--database` or `--workgroup` where either was renamed. This command always names a -workgroup, so a query it runs is always under a cutoff. Athena's `primary` workgroup has no cutoff -at all, and what lands there is a query sent by something else. The console, the AWS CLI, and a -script of your own that left the workgroup out all do. - -Credentials and profile come from the AWS SDK's default chain, the same one the AWS CLI reads. -There is nothing Rainlytics-specific to configure. - -## The region has to be the one the data is in +Credentials and default region come from the AWS SDK credential chain. The selected region must +contain the workgroup and Glue table. Keep the table in the log bucket's region to avoid +cross-region S3 transfer on every query. -The region comes from that chain too. It reads `AWS_REGION` first and then the region on the -profile, and `--region` names one over the top of both: +## Always restrict partitions -```bash -rainlytics query "SELECT count(*) FROM cloudfront_logs" --region us-east-1 -``` +Athena charges by bytes scanned. For an hourly table, restrict `distributionid`, `year`, `month`, +`day` and `hour` where possible. -A workgroup, a Glue table and an S3 bucket each exist in one region. Ask a region that holds none of -them and Athena answers about the workgroup, which is the first thing it looks for: - -```text -rainlytics: WorkGroup rainlytics is not found. Athena was asked in eu-west-2. Name another with ---region. +```sql +SELECT count(*) +FROM cloudfront_logs +WHERE distributionid = 'E1EXAMPLE1234' + AND year = '2026' + AND month = '09' + AND day = '01' + AND hour IN ('13', '14') ``` -The first sentence is Athena's, and it names the workgroup it could not find and never says where it -looked. The second is this command adding that. A profile defaulting to a region the deployment -never went near produces exactly this, and so does a workgroup that really was deleted. - -Which region to ask is decided by the log bucket. Athena reads a bucket in another region and bills -the transfer for every query, which is why the [log table](../log-table/) belongs in the bucket's -region as well. Log delivery is the one part of Rainlytics that has to be configured from us-east-1. -Where that delivery stack sits is a separate question from where a query runs. - -Every command that reaches Athena takes `--region`, `--database` and `--workgroup`, including the -four [named questions](../rollups/). - -## Naming a partition is most of what a query costs +A condition on `timestamp_ms`, path, status or another normal column filters rows after Athena has +read the partition. Add timestamp bounds for exact edges, but keep the partition conditions. -Athena bills per byte scanned. The partition keys are `distributionid`, `year`, `month`, `day` and -`hour`, and a predicate on any of them cuts what is read before it is read. A predicate on anything -else narrows the rows afterwards, and the bytes are billed either way. - -Here is the difference on a real bucket. Figures read off the Chinese Boost log bucket on -2026-08-27, two days into delivery: - -```bash -rainlytics query "SELECT count(*) FROM cloudfront_logs" -``` +The workgroup stops a query that passes its byte limit. Athena charges for bytes read before a +cancelled query stops. Narrow the partitions or raise `bytesScannedCutoff` when a legitimate report +needs more data. -```text -Scanned 8.12 MB in 1.2s, billed as 10.0 MB (the per-query minimum). About $0.000050 at the -us-east-1 rate. -``` +## Read CloudFront values -```bash -rainlytics query "SELECT count(*) FROM cloudfront_logs - WHERE year = '2026' AND month = '08' AND day = '26' AND hour = '14'" -``` +Every column is stored as text. Cast numbers and timestamps in SQL. CloudFront uses `-` for missing +values and percent-encodes logged fields. -```text -Scanned 265 KB in 0.4s, billed as 10.0 MB (the per-query minimum). About $0.000050 at the -us-east-1 rate. +```sql +SELECT + from_unixtime(cast(timestamp_ms AS bigint) / 1000) AS requested_at, + cast(sc_status AS integer) AS status, + nullif(cs_referer, '-') AS referrer, + url_decode(cs_uri_stem) AS path +FROM cloudfront_logs +WHERE year = '2026' AND month = '09' AND day = '01' ``` -580 objects against 13. The second query read a thirtieth of what the first did, and both cost the -same, because Athena bills a ten million byte minimum whatever a query reads. - -That is the honest version of this example, and it is worth sitting with. Pruning saves nothing on -a two-day-old dataset. What it changes is the shape of the curve. The hour query goes on reading 265 -KB for ever, while the unqualified one grows with the bucket. [#9](https://github.com/KensioSoftware/rainlytics/issues/9) -measured this site levelling off near 1.6 GB under the 370-day expiry, at which point the same pair -reads 1.6 GB against 265 KB and costs $0.0080 against $0.000050. - -A rollup that runs every hour for a year is where that difference stops being academic. +See [Log table](../log-table/#column-names-and-values) for field-name mapping. -## The price is on standard error +## Cost and output -Every query reports what it scanned and what that came to: - -```text -Query 8d0a2f4c-1a3e-4f77-9d0e-6c2b1f9a4e11 ran in workgroup rainlytics in us-east-1. -Scanned 265 KB in 0.4s, billed as 10.0 MB (the per-query minimum). About $0.000050 at the -us-east-1 rate. -``` - -It goes to standard error. A pipeline reads rows, and a person still sees the price: +After each query, standard error reports the query ID, workgroup, region, bytes scanned, duration +and estimated cost. Result rows go to standard output. ```bash -rainlytics query "SELECT c_country, count(*) FROM cloudfront_logs - WHERE year = '2026' AND month = '08' +rainlytics query "SELECT c_country, count(*) AS views + FROM cloudfront_logs + WHERE year = '2026' AND month = '09' GROUP BY 1" | jq '.[0]' ``` -The dollar figure is an estimate, and the line says which rate it used. It applies the us-east-1 -rate of $5.00 per terabyte and the ten million byte minimum, both read from the AWS Pricing API on -2026-08-27, and rounds a scan up to the next megabyte the way the pricing page describes. Every -other region charges its own rate, and the invoice rounds across a month of queries. +Use `--output json`, `csv` or `table`. Athena returns results in pages, and Rainlytics fetches every +page. Add a SQL `LIMIT` when the complete result is too large to be useful. -A query that failed is priced at nothing, because Athena bills nothing for one. What it read -before giving up is still reported: - -```text -Scanned 1.20 GB in 8.4s. Athena does not charge for a query that failed. -``` - -A query the workgroup stopped is a different case. Athena cancels that one rather than failing it, -and it bills a cancelled query for what it scanned. - -## When the workgroup stops a query - -The workgroup carries a ceiling. A query that would scan past it is stopped, and the message says -what it read, what the workgroup allows, and the two ways forward: - -```text -rainlytics: Bytes scanned limit was exceeded. The query scanned 12884901888 bytes, and -workgroup rainlytics allows 10737418240 per query. -That ceiling is the workgroup's, and it is there so one query cannot run up a bill nobody -chose. Narrow the query by naming distributionid, year, month, day or hour, or raise -bytesScannedCutoff on the rainlytics workgroup if the query really needs to read that much. -``` - -Athena bills a cancelled query for what it scanned on the way to being stopped, so the ceiling -puts a bound on the cost without removing it. The [query workgroup](../query-workgroup/) page has -where the default comes from and how to move it. +## Permissions -## Output +The caller needs Athena access to the workgroup, Glue reads for the catalog and table, S3 reads on +the log bucket, and S3 reads and writes on the results bucket. -`--output json`, `csv` or `table`, defaulting to a table at a terminal and to JSON when standard -output is piped or redirected. Every value comes back as a string, because every column in the log -table is a string. +Grant these permissions from CDK: -```bash -rainlytics query "SELECT cs_uri_stem FROM cloudfront_logs - WHERE year = '2026' AND month = '08' AND day = '27'" --output csv > paths.csv +```typescript +workgroup.grantQuerying(role, table); ``` -A result larger than one page is fetched whole, since Athena hands back a thousand rows at a time -and a truncated answer would look like a complete one. A query whose answer is too big to hold is -one that wanted a `LIMIT`. - -## Reading the data - -Every column in the table is a string. The casting belongs in the query: - -```sql -SELECT - from_unixtime(cast(timestamp_ms AS bigint) / 1000) AS at, - cast(sc_status AS integer) AS status, - nullif(cs_referer, '-') AS referrer, - url_decode(cs_uri_stem) AS path -FROM cloudfront_logs -WHERE year = '2026' AND month = '08' AND day = '27' -``` - -CloudFront URL-encodes what it logs and writes `-` for an empty field. The [log -table](../log-table/) page has the full column list and where each name comes from. - -## Exit codes - -`0` when the query ran, `1` when it ran and could not finish, and `2` when the command line could -not be read. A query Athena refuses exits `1` with its reason on standard error, and a query that -was never sent, such as one whose SQL the shell took apart, exits `2`. - -## Permissions - -Held by whoever runs the query rather than by the deploy role. The [query -workgroup](../query-workgroup/) page lists them: Athena on the workgroup, Glue on the table and its -catalog, S3 read on the log bucket and S3 read and write on the results bucket. - diff --git a/docs/reports/README.md b/docs/reports/README.md index a814f7c..f12007f 100644 --- a/docs/reports/README.md +++ b/docs/reports/README.md @@ -1,9 +1,7 @@ # Calendar reports -`RollupSummaries` precomputes one JSON report for every closed day, week, month and year. EventBridge -Scheduler invokes a separate report Lambda once a day. The function composes stored summaries where -their arithmetic is safe, runs a period-wide Athena query where it is not, and writes the finished -document to the summaries bucket. +`RollupSummaries` writes JSON reports for closed days, weeks, months and years. Each report contains +several analytics sections for one calendar period. ```typescript new RollupSummaries(this, "Summaries", { @@ -14,27 +12,26 @@ new RollupSummaries(this, "Summaries", { }); ``` -The defaults use UTC and Monday. The report job runs 30 minutes after local midnight, after the -default summary run at 15 minutes past. It recomputes reports for the two most recently closed days. -A day that also closes a week, month or year causes that larger period to be written too. +The defaults use UTC and Monday. A daily schedule runs 30 minutes after local midnight and +recomputes the two most recently closed days. It also writes a week, month or year when that period +has just closed. -The recomputation is intentional. CloudFront can deliver a late object after the first summary run. -The next run rebuilds that summary, and the report writer reads it again. The existing report is -never an input. A successful rerun replaces the object at the same deterministic key. +## Read a report -## The document +```bash +rainlytics report day 2026-08-30 +rainlytics report week 2026-08-24 --time-zone Europe/London +rainlytics report month 2026-07 +rainlytics report year 2025 +``` -`ReportDocument` describes each stored document. `reportPeriod`, `reportSection`, `reportDocument` -and `reportKey` are also exported for code that reads or builds the schema. +The date selects a period. A weekly date can be any date inside the week. `--time-zone` and +`--week-starts-on` must match the deployment because both values are part of the S3 key. -```typescript -import { - reportDocument, - reportKey, - reportPeriod, - reportSection, -} from "@kensio/rainlytics"; -``` +The command writes the complete JSON report to standard output. Bucket, key, age and S3 request cost +go to standard error. A report read never starts Athena. + +## Document format ```json { @@ -56,13 +53,7 @@ import { "computedAt": "2026-08-30T23:30:03.001Z", "sections": [ { - "question": { - "name": "status-codes", - "includeBots": false, - "limit": 20, - "param": "q", - "redirectStatuses": ["302", "303", "307"] - }, + "question": { "name": "pageviews", "includeBots": false }, "accuracy": "approximate", "composition": "ranked-summaries", "source": { @@ -73,126 +64,49 @@ import { }, "value": { "type": "rows", - "columns": ["status", "responses"], - "rows": [{ "status": "200", "responses": "18492" }] - } - }, - { - "question": { - "name": "web-vitals", - "includeBots": false, - "limit": 20, - "param": "q", - "redirectStatuses": ["302", "303", "307"] - }, - "accuracy": "exact", - "composition": "period-query", - "source": { - "from": "2026-08-23T23:00:00.000Z", - "until": "2026-08-30T23:00:00.000Z", - "summaries": 0, - "queries": 1, - "complete": true - }, - "value": { - "type": "rows", - "columns": ["metric", "p75"], - "rows": [{ "metric": "LCP", "p75": "2450" }] + "columns": ["path", "views"], + "rows": [{ "path": "/", "views": "18492" }] } } ] } ``` -`period` records the local calendar dates and their UTC instants. `sourceCoverage` is the outer span -represented by the section sources. Its `complete` field is true when at least one source covers the -whole report period without a gap. A document whose expected rollups are all missing carries -`null`. - -`computedAt` is the instant the document was assembled. The builder refuses an instant before the -period closes. - -Every section records its `SummaryQuestion`. A filter or limit that changes the stored answer stays -attached to the value in the report. The `source` records the span and the number of summaries or -period queries used. - -## Calendar boundaries - -`reportPeriod` accepts any instant inside a day, week, month or year. An IANA time zone supplies the -calendar. Days begin at local midnight. Months begin on their first local date and years begin on 1 -January. - -Weeks begin on Monday by default. Passing `weekStartsOn` selects another weekday, and a weekly -period records the choice. The other units omit it because it has no effect on their boundaries. +`period` carries local dates and the exact UTC range. A daylight-saving change can make a local day +23 or 25 hours long. -The UTC duration follows the local calendar. A day across a daylight-saving change can contain 23 or -25 hours. A week containing that day changes length with it. `from` and `until` record the resulting -UTC instants. `startsOn` and `endsBefore` record the local dates. +Each section records the question, calculation method, source coverage, accuracy and value. A +missing or malformed source produces an unavailable section rather than a partial value presented +as complete. -Reports cover closed periods. The second argument to `reportPeriod` is the computation clock. The -builder accepts the period when `until` is equal to or earlier than that clock. It raises a -`RangeError` while the current period is open. +Import the builders and types from the package root: ```typescript -const period = reportPeriod( - { - unit: "month", - at: new Date("2026-07-15T12:00:00Z"), - timeZone: "Europe/London", - }, - new Date("2026-08-01T00:00:00Z"), -); +import { + reportDocument, + reportKey, + reportPeriod, + reportSection, + type ReportDocument, +} from "@kensio/rainlytics"; ``` -The clock defaults to the current instant where a caller leaves it out. The scheduled writer passes -its invocation time so the decision is reproducible. - ## How sections are calculated -The writer uses stored summaries when a rollup exposes serialisable addition rules. It chooses the -largest available UTC windows that exactly cover the report. A UTC report normally reads daily -summaries. A time zone offset from UTC uses hourly summaries at its edges and daily summaries in its -interior. - -Additive counts remain exact. A percentage remains exact when the stored totals expose the counts -needed to recompute it. Ranked answers become approximate when they combine several summaries. Each -summary has already discarded rows below its limit, so a full-period ranking cannot recover every -candidate. - -The writer runs one Athena query over the whole period where summaries cannot produce the right -answer. Percentiles use this path. So does a derived value whose recomputation is only available as -JavaScript, including the default cache hit ratio. The section records `period-query`, zero -summaries and one query. Its result is exact. +The report job combines stored summaries when their values can be combined correctly. Additive +counts remain exact. Rankings built from several truncated summaries are marked approximate. -Visitor counts always use a period query. Daily summary salts deliberately prevent identities from -linking across days. The report query derives a separate salt for the whole calendar period, which -allows one person to count once without reusing an identifier from another period. +The job runs a period-wide Athena query when summary values cannot reproduce the answer. This path +is used for percentiles, period visitor counts and derived values whose required raw totals are not +stored. Those sections are marked `period-query`. -`reportSection` applies the following rules when code builds a section directly from summaries. +Visitor counts use one salt derived for the complete report period. This lets a browser count once +without linking its identifier to another calendar period. -| Rule | One summary spanning the report | Several summaries covering the report | -| --------------- | ------------------------------- | ------------------------------------- | -| `additive` | exact | exact | -| `ranked` | exact | approximate | -| `visitor-count` | exact | unavailable | -| `percentile` | exact | unavailable | +The report writer does not use an Athena query to hide missing summary windows. A gap makes the +affected section unavailable with `incomplete-source`. -The scheduled writer avoids the last two unavailable results by using a period query. - -## Incomplete sources - -A missing, malformed or mismatched summary becomes a gap. The writer does not hide such a gap with -a raw query. The affected section has `accuracy` set to `unavailable`, `reason` set to -`incomplete-source`, and `value` set to `null`. Its source metadata covers only the summaries that -were actually read, so the document cannot present incomplete data as a complete answer. - -An optional rollup that was expected but never stored can be represented as `missing-rollup` with a -null source. This keeps a report containing no measurements distinct from one whose writer never -looked for the question. - -## S3 keys and retention - -`reportKey` derives the key from the schema version and period. +## Object keys ```text reports/v1/UTC/day/2026-08-24.json @@ -200,132 +114,39 @@ reports/v1/Europe%2FLondon/week/monday/2026-08-24.json reports/v1/Asia%2FTokyo/month/2026-08-01.json ``` -The escaped time zone occupies one path segment. A weekly key includes its first weekday. The local -opening date sorts periods in calendar order under the prefix. A rerun sends another S3 `PutObject` -to this key, so readers see the recomputed document without finding a new location. +A rerun writes the same key. Readers then see corrections for late logs without discovering a new +object. The schema version appears in the key and the document. -The schema version appears in the key and document. A reader asks for the version it understands. -Changing a field's meaning or removing it requires a new version. An optional field can be added to -the current version. +Raw log retention must cover the longest report period you need to recompute. The default 370 days +covers an annual report and its next scheduled rebuild. -The default raw log retention is 370 days. That covers a 366-day annual report and leaves four days -for its scheduled recomputation. Shortening `LogBucket.retention` below the largest report period -can make that report unavailable. The report documents themselves remain in the summaries bucket. - -## Reading a report - -The command line selects a report from its calendar period and reads the document with one S3 GET. -It derives the key from the unit, date, time zone and first weekday. - -```bash -rainlytics report day 2026-08-30 --summaries rainlytics-summaries-1a2b -rainlytics report week 2026-08-24 --time-zone Europe/London -rainlytics report month 2026-07 -rainlytics report year 2025 -``` - -`--time-zone` must match the `RollupSummaries` deployment. For a weekly report, -`--week-starts-on` must also match. The defaults are UTC and Monday. The bucket comes from -`--summaries` or `RAINLYTICS_SUMMARY_BUCKET`, and the region comes from `--region` or the AWS SDK's -default chain. - -The versioned document is written unchanged as JSON on standard output. The bucket, object key, -object age and one-GET cost go to standard error. A missing, incomplete or unsupported document -leaves standard output empty and exits non-zero. The reader never runs Athena. - -## Comparing adjacent periods - -`reportComparison` compares a closed report with the immediately preceding period of the same -calendar unit. It uses the time zone recorded by the current report and the first weekday recorded -for a weekly report. Calendar arithmetic selects the earlier period, including weeks with a -configured first day and periods around daylight-saving changes. - -```typescript -import { previousReportPeriod, reportComparison } from "@kensio/rainlytics"; - -const previousPeriod = previousReportPeriod(current.period); -const previous = await loadReport(previousPeriod); -const comparison = reportComparison({ current, previous }); -``` - -The comparison is a derived result with its own schema version. Stored report documents remain the -source. The report writer can overwrite either document when late logs arrive. A stored comparison -would then describe an older pair of values until another job refreshed it. Deriving the result -from both documents keeps recomputation in one place and needs no Athena query. - -Pass `--compare` to ask the command line for the derived result: +## Compare adjacent periods ```bash rainlytics report month 2026-07 --compare ``` -The command reads the selected report first. It then reads the preceding report with one additional -S3 GET. Standard output remains one JSON document. Standard error names both keys, their ages and -the cost of two GET requests. - -The comparison carries document metadata for both reports. Each available section also carries the -two section sources, their calculation methods and a combined accuracy. The combined accuracy is -`approximate` when either source section is approximate. - -Metric changes follow these rules: +The command reads the selected report and the immediately preceding report. It calculates a +versioned comparison document without Athena. -| Metric | Change | Unit and direction | -| ---------------------------------------- | ------------------- | ----------------------------------------------------------------------------------------- | -| Counts, including pageviews and visitors | Relative percentage | The metric's count unit. Movement is unrated. | -| Cache hit ratio | Percentage points | Percent. A higher value is better. | -| Web Vitals p75 | Relative percentage | Milliseconds for LCP, FCP and TTFB. CLS uses its unitless score. A lower value is better. | -| Caller-defined durations | Relative percentage | Supplied by the metric definition. The definition also supplies the preferred direction. | +Counts use relative percentage change. Cache hit ratio uses percentage points. Web Vital values use +relative percentage change and treat lower values as better. A zero baseline produces a `null` +relative change with a `zero-baseline` reason rather than infinity. -A zero baseline produces a `null` relative change with `reason` set to `zero-baseline`. The current -value, previous value, numeric difference and trend remain available. JSON never contains infinity. - -Rows are matched on every non-metric column. A row present on one side only has an unavailable -comparison. Ranked questions use `ranked-row-absent` as the reason because the row may sit below the -other period's stored limit. Its missing value stays `null` and never becomes zero. - -Sections are withheld when a source is incomplete or unavailable, the question configuration -changed, the columns changed, or either report lacks the section. Rainlytics supplies definitions -for its shipped questions. A caller can pass `definitions` to `reportComparison` for a custom -question. A custom definition names the numeric columns, their units, the change measure and the -preferred direction. +Ranked rows are matched by their non-metric columns. A row present on one side only is unavailable +because it may have fallen below the other report's stored limit. ## Cost -The report path has no reserved or hourly capacity. Its AWS services charge for invocations, -duration, requests, bytes stored and bytes scanned. - -The default deployment creates one Scheduler invocation and one 512 MB Lambda invocation each day. -AWS includes 14 million Scheduler invocations per month and one million Lambda requests plus 400,000 -GB-seconds per month in their free tiers. The report schedule is 30 invocations per month. - -The six default rollups issue two period queries per report. One computes the cache hit ratio and one -counts pageview visitors. Recomputing two closing days writes about 72 reports and runs about 143 -queries in an average month. At Athena's 10 MB minimum and $5 per TB, those queries cost about -$0.0072. Actual cost rises when a period query scans more than 10 MB. - -For UTC reports, composing the other five questions reads about 1,217 summary objects and writes -about 72 report objects per month. At the US East (N. Virginia) S3 Standard request prices of $0.0004 -per 1,000 GET requests and $0.005 per 1,000 PUT requests, those requests cost about $0.00085. The -small JSON documents add a fraction of a cent in storage. Other Regions can have different prices. +The report path uses Scheduler, Lambda, Athena and S3 on demand. It reserves no capacity. -The report writer therefore adds about one cent per month for a quiet default deployment whose -Lambda usage stays inside the free tier and whose Athena queries stay at the minimum. This estimate -does not include the existing summary jobs. Traffic volume and added period-query rollups determine -the variable part. - -Prices were checked on 31 August 2026 against the [Athena pricing page](https://aws.amazon.com/athena/pricing/), -[EventBridge pricing page](https://aws.amazon.com/eventbridge/pricing/), [Lambda pricing -page](https://aws.amazon.com/lambda/pricing/) and [S3 pricing page](https://aws.amazon.com/s3/pricing/). +Most report sections reuse stored summaries. Period-wide Athena queries determine the variable +part of the cost. A quiet default deployment whose queries stay at Athena's minimum adds roughly a +cent a month for reports, excluding the summary jobs themselves. Traffic, optional questions and +regional prices change that estimate. diff --git a/docs/rollups/README.md b/docs/rollups/README.md index cbeb66e..5119d96 100644 --- a/docs/rollups/README.md +++ b/docs/rollups/README.md @@ -1,370 +1,123 @@ # Rollups -The questions the command line answers without anybody writing SQL. +A rollup is a named analytics question. Rainlytics generates its Athena SQL, schedules it and gives +it a command-line name. ```bash rainlytics pageviews --last 7d ``` -```text -path views ------------ ----- -/ 412 -/liju/ 208 -/grammar/ 97 -``` - -`referrers`, `browsers`, `status-codes`, `cache-hit-ratio` and [`searches`](../searches/) are the -other default questions. [`javascript-errors`](../javascript-errors/) and -[`web-vitals`](../web-vitals/) are shipped commands for a deployment using the optional browser -beacon. Each takes the same `--last`, the same `--path` and `--host`, the same output formats and the -same bot filter. - -Each of them answers from the [precomputed summaries](../summaries/) a schedule wrote, at the cost of -a GET per window. `--query` runs the question through Athena for a fresher answer, and reports what -it scanned and what that cost. [Reading a precomputed -answer](#reading-a-precomputed-answer) below has when each applies. - -## What each one counts - -**`pageviews`** counts the pages people looked at. A pageview is a GET that answered HTML and -succeeded, which is what separates a page from the images, stylesheets and fonts the same log -records. A 304 counts, because a browser being told its copy is current is somebody looking at the -page. The path is decoded, for the reason under [The log is percent-encoded -twice](#the-log-is-percent-encoded-twice) below. - -**`referrers`** counts where people arrived from, by host. Requests carrying no referrer are left -out, and so are the ones this site sent itself. Those are somebody moving around inside it. On the -reference site an unfiltered version of this is topped by its own stylesheet. - -**`browsers`** counts the same pageviews by browser family and device class. A browser is Edge, -Opera, Samsung Internet, Firefox, Chrome family, Safari family or Other. A device is Tablet, Mobile, -Desktop or Other. The answer groups on both, so Chrome family on Mobile and Chrome family on Desktop -are separate rows. - -The classifier is a fixed `CASE` ladder in the Athena query. Specific Chromium browsers are checked -before the Chrome and Safari compatibility tokens their user agents also carry. This keeps the -scheduled summary and `rainlytics browsers --query` on one definition, with no parser or second pass -in Lambda. New browsers fall into the family they present as or into Other. Existing raw rows are -classified by the current query whenever they are read again. - -The user agent cannot support a precise answer. Chromium derivatives can present the same string as -Chrome. iPadOS can present as macOS. Android without a Mobile token is called a tablet, which can -also catch a television or another embedded device. Unusual and reduced agents can land under the -wrong class or Other, and a client can write any user agent it likes. Browser versions are left out. -This rollup reads the access log only and uses nothing reported by the optional beacon. - -**`status-codes`** counts every response, including the assets the pageview count leaves out. A -stylesheet returning 404 is worth seeing and a rollup looking only at pages never would. Requests to -the [beacon's own path](#the-beacons-own-requests) are the exception. - -**`cache-hit-ratio`** counts over the values that say whether the cache served the request, being a -Hit, a RefreshHit and a Miss. A redirect and the `FunctionGeneratedResponse` a beacon event comes -back as never reached the cache, so counting them would move the ratio without the cache having -changed. An `Error` is left out for a different reason: CloudFront caches error responses, and -`Error` also covers a viewer who disconnected after being served one, so it says too little either -way. - -Three rollups wait for a deployment to opt in. [`javascript-errors`](../javascript-errors/) counts -uncaught exceptions and unhandled rejections by page and message. [`web-vitals`](../web-vitals/) -calculates p75 for each reported vital. Both have their own commands. -[`beacon-events`](../beacon-events/) counts every event and runs through `saved-query`. A site with -no beacon then pays for none of those empty answers. - -## The log is percent-encoded twice - -CloudFront percent-encodes every value it writes into a log record, and a request URI reaches it -already carrying the browser's own encoding. A page at `/words/好/` is requested as -`/words/%E5%A5%BD/` and recorded as `/words/%25E5%25A5%25BD/`. - -`pageviews` decodes the path twice, so it reports the address a reader would recognise. One pass -answers `/words/%E5%A5%BD/`. That is the URI the browser sent, and it reads no better than the -record. A site whose addresses are all ASCII sees the same table either way. - -Only `pageviews` reads a column carrying the encoding. `referrers` reads a referrer for its host, -and a host is ASCII whatever the rest of the URL holds. A status code and a result type arrive -plain. The crawler filter matches ASCII substrings of a user agent and reads an encoded one the -same way. - -Two limits are worth knowing. `url_decode` reads `+` as a space. That is right for a query string -and wrong for a path, where `+` is a literal. Athena also raises over an escape naming no byte, -such as `%zz`. A path carrying one that still answered HTML would fail the query outright. Both -stayed theoretical across 137,000 records of real traffic. - -## Crawlers are most of the traffic - -Every rollup leaves automated traffic out by default. That is a judgement, and here it is. - -One hour of the reference site in August 2026 held 9,492 requests. 3,748 of them matched the bot -filter, and 1,951 were a single crawler. Bots were 39% of the hour and the largest single user agent -on the site. An unfiltered pageview count is not the raw number with the opinions taken out. It is a -number that says more about crawlers than about anybody who reads the site. - -The filter matches four substrings against a lowercased `cs(User-Agent)`: - -```text -bot|crawl|spider|slurp -``` +## Built-in questions -Substrings, because a crawler names itself `ClaudeBot/1.0` with the token glued to a word, which -a whole-word match would walk past. The cost is a device whose name happens to contain one, and the -Cubot range of Android phones is the example. Count it both ways to see how much of the difference is yours: - -```bash -rainlytics pageviews --last 7d --include-bots -``` +| Command | Answer | +| ----------------- | --------------------------------------------------------------------- | +| `pageviews` | Successful HTML GET requests by decoded path. A 304 counts as a view. | +| `referrers` | External referrer hosts. Empty and same-site referrers are omitted. | +| `browsers` | Pageviews by browser family and device class. | +| `status-codes` | All site responses by HTTP status, excluding the beacon path. | +| `cache-hit-ratio` | Hits and misses for requests that reached the cache. | +| `searches` | Search terms and temporary redirects from configured search pages. | -`status-codes` is the one where `--include-bots` is usually what you want. Bots find the broken -links first and in numbers. +The exported `rollups` array contains these six questions. -## The beacon's own requests +`javascript-errors` and `web-vitals` are optional questions for sites using the browser module. +`beacon-events` is another optional rollup and runs through `saved-query`. Optional questions are +excluded from the defaults because a site without those browser events would pay for empty queries. -The beacon sends a GET to `/_rainlytics` on the site's own domain and carries its payload in the -query string. An event is another row in the same log. It writes one row per event, and a -single-page app reporting route changes, web vitals and errors sends several per reader per page. A -quiet site can end up with more beacon rows than responses of its own. +## Filter a question -`status-codes` leaves those requests out, and no option puts them back. Every event answers 204. A -window of them leads the table under one status, and the 404 the question exists to surface sits -somewhere below it. The question is what the site answered for the things it serves, and the beacon -is Rainlytics measuring the site. - -Anybody checking that the beacon is delivering has [`rainlytics query`](../query/): +All named questions support a time range: ```bash -rainlytics query "SELECT sc_status, count(*) AS events FROM cloudfront_logs - WHERE year = '2026' AND month = '08' AND day = '29' - AND strpos(cs_uri_stem, '/_rainlytics') = 1 - GROUP BY 1" +rainlytics pageviews --last 24h +rainlytics referrers --last 2w ``` -The path is one constant, `defaultBeaconPath`, that the beacon and the rollup both read. Naming a -different one belongs to the beacon construct. - -The other five questions leave beacon rows out already, for reasons they had anyway: - -- **`pageviews`**, **`referrers`** and **`browsers`** count a GET that answered `text/html` with a - 200 or a 304. An event answers 204 and names no content type. -- **`searches`** wants its parameter non-empty. A payload names its parameters `v`, `e` and `p`, and - carries no `q`. -- **`cache-hit-ratio`** counts a Hit, a RefreshHit or a Miss. A CloudFront Function answers every - event, and the cache is never asked. +The suffix can be `h`, `d` or `w`. The range becomes partition predicates, which limit the bytes +Athena reads. -Those five are checked against delivered beacon records rather than taken on trust, in -`src/beacon-events.test.ts`. Each of them leaves the rows out through a condition it has for its own -reasons, and a rollup of your own gets none of that for free. - -## `--path` and `--host` narrow the question - -`--path` counts one section of a site, as a prefix of the address: +Filter by path prefix or host: ```bash rainlytics pageviews --path /guides/ --last 30d +rainlytics status-codes --host docs.example.com --last 7d ``` -It matches the address a reader sees. The record holds it percent-encoded twice, and the filter -decodes before comparing, so `--path /词典/` finds the pages `pageviews` prints under that name. The -text is taken literally. A path holding `_` or `%` matches itself. +Repeat `--path` to combine several sections. Host matches are exact. -Give `--path` again for each section that belongs in one answer. A request counts when its address -starts with any of them: +Automated traffic is omitted by default. The filter matches `bot`, `crawl`, `spider` or `slurp` in +the lowercased user agent. Include those requests when they matter to the question: ```bash -rainlytics pageviews --path /guides/ --path /tutorials/ --last 7d +rainlytics status-codes --last 7d --include-bots ``` -Guides and tutorials are one section of a site to whoever writes them, and a site with search boxes -at `/words/search/` and `/sentences/search/` has no prefix covering both. Each path becomes its own -prefix test and the tests are joined by `OR`. One `--path` writes what it always wrote. +The filter is useful but cannot identify every automated client. Any client controls its own user +agent. -An answer counting several sections together says nothing about which of them a row came from. -[`searches`](../searches/) names the section on every row when it is given more than one, and a -rollup of your own gets the same column from [`matchedPath`](#naming-the-section-a-row-came-from). +## Read stored or fresh data -`--host` counts one of the sites a single distribution serves: +Named commands read [rollup summaries](../summaries/) from S3: ```bash -rainlytics status-codes --host docs.example.com --last 7d +rainlytics pageviews --last 7d --summaries rainlytics-summaries-1a2b ``` -That one matches in full. A site and its `www` name are two hosts, and folding them together is a -decision for whoever runs them rather than a default. `x-host-header` is in the delivered field set -for exactly this, and nothing else in a record says which site was asked for. - -Neither option changes what a query costs. `--last` has already decided which partitions are read, -and these two narrow rows that are paid for either way. Narrowing to one section of a busy site -gives a shorter answer for the same money. - -## `--last` decides what the question costs +The command combines complete hourly and daily windows inside the requested range. The current +partial hour is left out. Standard error reports the exact span used, missing edge windows and the +age of the newest summary. -A range becomes partition predicates rather than a filter on the record's own timestamp. `--last 7d` -over a year of logs reads seven days of objects. The same range written `WHERE timestamp_ms > ...` -answers identically and reads the year to do it, which is the mistake the whole partition layout -exists to prevent. +Add `--query` to calculate one fresh result from raw logs: ```bash -rainlytics pageviews --last 24h -rainlytics referrers --last 2w -rainlytics status-codes --last 4w --include-bots +rainlytics pageviews --last 7d --query ``` -Whole hours, days or weeks. There is no month, because a month is not a fixed length and a range -that quietly meant thirty days would be worse than one nobody could ask for. - -The predicate a range builds names each partition key separately: - -```sql -WHERE year IN ('2026') - AND month IN ('08', '09') - AND day IN ('28', '29', '30', '31', '01', '02', '03') - AND cast(timestamp_ms AS bigint) BETWEEN 1787875200000 AND 1788436800000 -``` +Stored ranked results are approximate across several windows because each window only kept its own +leading rows. Counts are added and the combined rows are ranked again. A fresh Athena query ranks +the full range in one pass. -Those are a cross product. A week spanning a month boundary asks for seven days in two months and -reads fourteen partitions, and the timestamp condition after them is what keeps the answer exact. The -alternative is one predicate per day joined by `OR`, which reads exactly the seven and which Athena -plans more slowly the longer the range gets. Fourteen partitions against a year of them is still the -difference the layout exists to make. +Percentiles and visitor identities cannot be combined from summary values. Commands that need the +raw distribution or identity set require a single stored window or `--query` for a larger range. -## The same SQL, saved in the console +## Save the generated SQL -The `RollupQueries` construct saves each rollup as an Athena named query, written by the same builder -the command writes with: +`RollupQueries` stores one Athena named query per rollup: ```typescript import { RollupQueries } from "@kensio/rainlytics/cdk"; -new RollupQueries(this, "RainlyticsRollups", { table, workgroup }); -``` - -Somebody in the console can then read what `rainlytics pageviews` counts, run it, and edit it into a -question of their own without reading this repository. - -The saved copies cover the current month. There is no span to compute at deploy time, and dates -baked in then would be the dates of whoever last deployed and would change the template on every -deploy. They ask Athena what month it is: - -```sql -WHERE year = date_format(current_date, '%Y') - AND month = date_format(current_date, '%m') -``` - -### Narrowing a saved copy - -Everything else a command takes is settled per rollup, by `requests`: - -```typescript -new RollupQueries(this, "RainlyticsRollups", { +new RollupQueries(this, "SavedQueries", { table, workgroup, requests: { - searches: { paths: ["/search/"], param: "term" }, + searches: { paths: ["/search/"], param: "q" }, }, }); ``` -`searches` is why this is here. It reads one query-string parameter on one page. A saved copy left -to the defaults counts every query string on the distribution, while its own description tells the -reader to name the search page with `--path`. The parameter defaults to `q`. A site whose box calls -it something else gets a saved query that answers with an empty table. - -Per rollup, and not one set of options across all five. `/search/` is the search page to `searches` -and one directory of a site to `pageviews`. A shared set would save `rainlytics-pageviews` as a -query counting the search page under a name promising the whole site. That is the same fault the -other way round. A rollup left out of `requests` takes the defaults a command starts from. - -An entry takes what a rollup command takes, apart from `--last`. The range is always the current -month for the reason above, and the database comes from the table: - -```typescript -requests: { - "status-codes": { includeBots: true }, - searches: { - host: "docs.example.com", - paths: ["/search/"], - param: "term", - redirectStatuses: ["301", "302"], - }, -} -``` - -`paths` is the list `--path` collects when a command is given it more than once. A site with a -search box under two sections names both, and the saved copy counts them together. - -`redirectStatuses` is `--redirect-status`, and it is where a site says what its own search page -answers with. The three a search counts by default are 302, 303 and 307, and -[searches](../searches/#which-statuses-count) covers why 301 and 308 are left out. A site whose -exact match answers 301 puts it here, and the saved query reads a `redirected` column that is right -for it. +Saved queries cover the current month. Run one with: -An entry takes whatever `RollupRequest` carries, minus the range and the dataset. A field added to -the request arrives here on its own. - -`RollupSummaries` takes the same shape on its own `requests` prop. A site narrowing a question for -both writes the narrowing once and passes the constant to each: - -```typescript -const searches = { paths: ["/liju/search/", "/cidian/search/"], param: "term" }; - -new RollupQueries(this, "RainlyticsRollups", { - table, - workgroup, - requests: { searches }, -}); -new RollupSummaries(this, "RainlyticsSummaries", { - table, - workgroup, - requests: { searches }, -}); -``` - -A saved query and a stored summary that drift apart answer two questions under one name. The -constant keeps them in step, and -[reading a precomputed answer](#a-summary-answers-the-question-it-was-computed-with) is where the -command line picks the same narrowing up. - -A fact that belongs to every question, such as the host of one site on a distribution serving -several, is a variable spread into each entry. Every key has to name a rollup being saved, and a -mistyped one fails at synthesis. The alternative is a deployed query still counting whatever it -counted before. - -Each saved description says what its own copy covers. The console shows the narrowing to somebody -who has read no SQL: - -```text -Count searches by the term somebody typed. What "rainlytics searches" runs. -Over the current month, on docs.example.com, under /search/, reading the "term" -parameter, counting 301 or 302 as redirected. +```bash +rainlytics saved-query searches ``` -The statuses are named there only where a deployment chose its own. The three a search counts by -default are in the rollup's own description already, and a line repeating them on every copy says -nothing about that copy. +Pass the same `requests` values to `RollupQueries` and `RollupSummaries` so the saved SQL and stored +answers describe the same question. -`--limit` is the one option left out of that line. A row count decides how much of the answer is -printed and leaves what was counted where it was. It sits on the last line of the SQL below. +## Write a custom rollup -## Writing a rollup of your own - -The rollups above are assembled from parts the package exports, and a site with a question of its -own assembles another the same way. A rollup is a name, some help text and a function that writes -the SQL for one request: +A custom rollup supplies a name, help text and a function that builds SQL for one request. ```typescript -import { - lastRange, - qualifiedTableName, - type Rollup, - rollupRequest, - rollupSql, - rowsFor, -} from "@kensio/rainlytics"; +import { qualifiedTableName, type Rollup, rowsFor } from "@kensio/rainlytics"; const countries: Rollup = { name: "countries", - summary: "Count views by country.", - description: "Counts where readers were, most read from first.", + summary: "Count pageviews by country.", + description: "Counts pageviews by viewer country, highest first.", isRanked: true, + totals: { added: ["views"] }, body: (request) => [ "SELECT c_country AS country, count(*) AS views", @@ -375,337 +128,50 @@ const countries: Rollup = { ` LIMIT ${String(request.limit)}`, ].join("\n"), }; - -const sql = rollupSql( - countries, - rollupRequest({ range: lastRange("7d", new Date()) }), -); -``` - -`rowsFor` writes the whole `WHERE` clause. The partition predicate, the timestamp bounds, the -crawler filter and the `host` and `paths` the request narrowed to all come out of it, and its second -argument carries the conditions this one question adds. Writing that by hand puts a second copy of -[what a range costs](#--last-decides-what-the-question-costs) and of [the crawler -filter](#crawlers-are-most-of-the-traffic) in the site's own repository, and the copy is the one -that goes stale. - -`rollupSql` hands back the text, and running it is the site's own Athena client. [Saving it in the -workgroup](#running-a-rollup-of-your-own) is the other way round, and the way that needs no client. - -`countsVisitors: true` puts a visitor count on the summaries a scheduled copy of the rollup writes. -The count is over pageviews under the same narrowing, whatever this question counts, and -[Counting visitors](../visitors/) has what it means and what it costs. - -### Reading a query-string parameter - -`decodedParameter` writes the expression that takes one parameter out of a record and decodes it. A -site counting the campaigns its inbound links name groups by that: - -```typescript -import { - decodedParameter, - qualifiedTableName, - type Rollup, - rowsFor, -} from "@kensio/rainlytics"; - -const campaign = decodedParameter("utm_campaign"); - -const campaigns: Rollup = { - name: "campaigns", - summary: "Count views by the campaign that sent them.", - description: "Counts the campaigns inbound links named, most sent first.", - isRanked: true, - body: (request) => - [ - `SELECT ${campaign} AS campaign, count(*) AS views`, - ` FROM ${qualifiedTableName(request.dataset)}`, - rowsFor(request, ["cs_uri_query <> '-'", `${campaign} <> ''`]), - " GROUP BY 1", - " ORDER BY 2 DESC, 1", - ` LIMIT ${String(request.limit)}`, - ].join("\n"), -}; -``` - -It names `cs_uri_stem` and `cs_uri_query` for itself. A record carries no whole URL, and those two -columns are joined back together with the `?` that was between them before CloudFront split them up. -(`'-'` is what CloudFront writes where a field was empty. The first condition drops the requests -that carried no query string.) - -The value comes back decoded once, where [pageviews](#the-log-is-percent-encoded-twice) decodes a -column twice. `url_extract_parameter` decodes its own answer and one further pass finishes the job. -A second pass would decode a term holding a percent sequence twice, and `50%` typed into a search -box is the case. That rule is what the function carries. A hand-written -`url_decode(url_extract_parameter(...))` in the site's own repository carries the expression and -leaves the rule behind. - -`decodedColumn` is the other half of this, for a question grouping by a whole column rather than by -one parameter. `pageviews` reads the path through it, and [`searches`](../searches/) reads its term -through `decodedParameter`. - -### Naming the section a row came from - -A question narrowed to several paths counts them together, and one term or one country then holds -rows from every one of them. `matchedPath` writes the prefix a row's address started with, as a -column the question selects and groups by: - -```typescript -import { - matchedPath, - qualifiedTableName, - type Rollup, - rowsFor, -} from "@kensio/rainlytics"; - -const countriesBySection: Rollup = { - name: "countries-by-section", - summary: "Count views by country and section.", - description: "Counts where readers were, section by section.", - isRanked: true, - body: (request) => - [ - `SELECT ${matchedPath(request)} AS section,`, - " c_country AS country, count(*) AS views", - ` FROM ${qualifiedTableName(request.dataset)}`, - rowsFor(request, ["sc_content_type LIKE 'text/html%'"]), - " GROUP BY 1, 2", - " ORDER BY 3 DESC, 1, 2", - ` LIMIT ${String(request.limit)}`, - ].join("\n"), -}; ``` -It is a `CASE` over the same prefix tests `rowsFor` filters with, branch by branch in the order the -request gave them. One definition of a prefix match covers both. A copy of the expression in a -site's own repository is a second definition, and the way those drift is a column that stops -agreeing with the filter beside it. +Use `rowsFor` for the `WHERE` clause. It adds the time partitions, timestamp bounds, bot filter, +host filter and path filters from the request. -What the column holds follows from how many paths a run was given: +`totals.added` lists numeric columns that can be added across stored windows. All other columns +identify a row. A rollup with no totals can only answer from one stored window. -- **Several.** The first one the address starts with. Where two overlap, `/guides/` given alongside - `/guides/advanced/` reports a row under the second as `/guides/`. Every row is then in exactly one - section, and a reader adding the rows up counts each of them once. -- **One.** That path, as a literal. Every row counted started with it, and a `CASE` there asks a - question with one answer. -- **None.** `CAST(NULL AS varchar)`. The whole distribution was counted and no prefix matched. An - empty string would claim a prefix nobody asked for, and the cast gives the column a type in the - result Athena hands back. - -A rollup selects it however many paths it was given, and `--path` decides what comes back. -[`searches`](../searches/) is the built-in one that reads it, for a site with two search boxes. - -The construct saves a site's rollup in the console beside the built-in ones: +Save and schedule the custom question: ```typescript import { rollups } from "@kensio/rainlytics"; -import { RollupQueries } from "@kensio/rainlytics/cdk"; -new RollupQueries(this, "RainlyticsRollups", { +const questions = [...rollups, countries]; + +new RollupQueries(this, "SavedQueries", { table, workgroup, - rollups: [...rollups, countries], + rollups: questions, }); -``` - -The saved copy covers the current month, as the built-in ones do. Its description says what it counts and -stops there, since there is no `rainlytics countries` to point a reader at. It takes an entry in -`requests` under its own name the way the built-in ones do. - -A name is lowercase words joined by hyphens (`cache-hit-ratio`). It becomes a CDK logical id and an -Athena query name, and `assertRollupName` refuses anything else at synthesis. - -### Adding a rollup of your own across windows - -A range of a week is 29 stored windows, and the command adds them together before it prints -anything. `totals` is where a rollup says how: - -```typescript -const countries: Rollup = { - name: "countries", - summary: "Count views by country.", - description: "Counts where readers were, most read from first.", - isRanked: true, - totals: { added: ["views"] }, - body: (request) => /* ... */, -}; -``` - -`added` names the columns holding counts. Every other column names a row, so two windows' rows are -matched on the country and their views add. The first count named is what a ranked answer is ordered -by, matching the `ORDER BY 2 DESC` the query writes for one window. - -A column worked out from the counts beside it is named under `recomputed`, and its function is handed -the counts of one row once they have been added: - -```typescript -totals: { - added: ["hits", "misses"], - recomputed: { - hit_percent: (added) => percentageOf(added["hits"], added["misses"]), - }, -}, -``` - -`cache-hit-ratio` is the shipped question that needs it. A percentage averaged across windows is a -figure about none of them, and the counts underneath it are what add. - -A rollup with no `totals` answers from one stored window. A range covering several is reported as -that, with `--query` offered for the span. That is the safe default for a question this package has -never seen, and a wrong guess would report a percentage as its own sum. - -### Running a rollup of your own - -The saved copy is what gives a site's own question a command line: - -```bash -rainlytics saved-query countries -``` - -`rainlytics saved-query` reads the queries saved in the workgroup and runs the one that matches, so -nothing on this side loads the site's code or asks for a build step. The name is the one Athena -lists, with or without the `rainlytics-` prefix, and a name matching nothing is answered with the -names that are saved there. -It takes `--output`, `--workgroup` and `--region`, and reports what the query scanned and what that -cost the way the built-in commands do. It takes no `--last` and no `--limit`. The SQL Athena holds -settled both when it was saved, which is why the range is the current month and the row count is the -one `requests` was given. The [command -line](../command-line/#running-a-query-saved-in-the-workgroup) page has the rest of it. - -### Replacing one of the built-in questions - -A site whose searches answer differently from the shipped `searches` writes its own version and -leaves the shipped one out of the list: - -```typescript -import { rollups } from "@kensio/rainlytics"; -import { RollupQueries } from "@kensio/rainlytics/cdk"; - -new RollupQueries(this, "RainlyticsRollups", { +new RollupSummaries(this, "Summaries", { table, workgroup, - rollups: [ - ...rollups.filter((rollup) => rollup.name !== "searches"), - mySearches, - ], + rollups: questions, }); ``` -`rainlytics-searches` in the console is then the site's own question, and `rainlytics saved-query -searches` runs it. The `rainlytics searches` command still runs the shipped one, since its command -list is the questions the package ships. Two ways of asking, and the saved one is the site's. - -Passing both is refused at synthesis, since one saved query cannot answer two questions: - -```text -More than one rollup is called "searches", and each would be saved as -"rainlytics-searches". Where one of them replaces a built-in question, leave -the built-in out: rollups: [...rollups.filter((rollup) => rollup.name !== -"searches"), mySearches] -``` - -## Output, and what happens next - -`--output json`, `csv` or `table`, defaulting to a table at a terminal and to JSON when piped. Every -value is a string, since every column in the log table is one. - -```bash -rainlytics pageviews --last 7d --output csv > pages.csv -rainlytics referrers --last 7d | jq '.[0].referrer' -``` - -`--limit` takes the top rows of a ranked rollup, twenty by default. `cache-hit-ratio` answers with -one row and has nothing to limit. - -## Reading a precomputed answer - -`rainlytics pageviews --last 7d` reads what the [summary schedule](../summary-schedule/) already -counted. The bucket comes from `--summaries` or from `RAINLYTICS_SUMMARY_BUCKET` in the environment, -and a range of a week is 29 objects and about a hundredth of a cent. - -```bash -rainlytics pageviews --last 7d --summaries rainlytics-summaries-1a2b -rainlytics pageviews --last 7d --query -``` - -The rows are the same either way. Standard error is where the two differ, and it carries the span -that answered and how old it is. - -### A summary answers the question it was computed with - -`--path`, `--host`, `--include-bots`, `--param` and `--redirect-status` each decide which requests -were counted, and a schedule cannot count every combination of them. `RollupSummaries` computes the -unfiltered form of each question, and [`requests`](../summary-schedule/) is where a deployment adds a -narrowed one under a name of its own. - -A run that names none of those five takes the ones the summaries were computed with. The -deployment declared its narrowing once and the command reads that copy back. A shell alias never has -to carry a second one. Standard error says which filters the run took: - -```text -Took --path /liju/search/ /cidian/search/ from the summaries. Those options -were left off this command line, and the answer covers the narrowing the -deployment computes. -``` - -An option somebody typed stays theirs. A run whose filters no stored summary matches is told what -was stored: - -```text -The stored pageviews summaries answer a different question. - --path: asked for /guides/, computed with the whole distribution -A schedule computes the questions its deployment named, and the requests prop -on RollupSummaries is where a narrowed one is added. --query answers this run -from Athena at the cost a query reports. -``` - -A change to `requests` leaves both narrowings in the bucket. Over a span holding some of each, the -command names the option that would settle it and stops: - -```text -The stored pageviews summaries over that span were not all computed the same -way, and this run named nothing to settle it with. - --path: some windows computed with /guides/, others with the whole distribution -``` - -Typing one of them settles nothing, since the windows computed the other way then refuse it. A span -on one side of the change reads from stored summaries, and `--query` answers one covering both. - -`--limit` is apart from those five. A row count decides how much of a ranked answer is printed and -leaves what was counted where it was, so a deployment computing the top hundred paths still answers -`rainlytics pageviews` with the top twenty. The stored hundred holds them. - -A summary computed with fewer rows than the command asks for is the other way round, and those rows -were never counted. A run that typed the count is refused. A run that typed none is cut to what the -stored windows hold, and standard error names the count it answered with. - -### Several windows add up, and the ranking is approximate - -A week is 29 stored windows and the command adds them together. Counts add. A row that fell outside -the stored rows of every window is missing from all of them, so a ranked answer assembled this way is -approximate and standard error says so. `--query` ranks the whole span in one pass. - -`cache-hit-ratio` adds its hits and its misses and works the percentage out again from the total. -Averaging two windows' percentages would answer a figure about neither. +Run it with `rainlytics saved-query countries`. The fixed CLI command list only contains questions +shipped by the package. -A visitor count belongs to one window and never adds. The identifier takes a new salt every day, so -two days' counts added together count everybody who came back twice over. A command reading several -windows says that it cannot give one. +Names use lowercase words joined by hyphens. Each scheduled question must have a unique name because +the name is part of its saved query and S3 key. -A rollup of your own says how its rows combine with [`totals`](#adding-a-rollup-of-your-own-across-windows), -and one that says nothing answers from a single stored window. +## Decode fields in custom rollups -## Anything else +Use `decodedColumn` for a whole logged field and `decodedParameter` for one query-string value. +`matchedPath` returns the path prefix matched by a request when a question covers several sections. -[`rainlytics query`](../query/) takes SQL. These are the questions worth a name, and the log table -holds a great many more. +These helpers keep custom questions consistent with the built-in URL decoding and path matching. diff --git a/docs/searches/README.md b/docs/searches/README.md index 1fa2ba6..c433666 100644 --- a/docs/searches/README.md +++ b/docs/searches/README.md @@ -1,6 +1,6 @@ # Searches -What people typed into a site's search box, counted off the access log. +`rainlytics searches` counts terms submitted to a search page from the CloudFront access log. ```bash rainlytics searches --path /search/ --last 30d @@ -9,112 +9,73 @@ rainlytics searches --path /search/ --last 30d ```text term searches redirected -------- -------- ---------- -好 41 38 -happy 12 0 +rain 41 38 +weather 12 0 ``` -What readers search for says which pages to write next, and a site usually has no record of it. The -terms are in the access log already. CloudFront records `cs-uri-query` whatever the cache key and -origin forwarding are set to, so a search answered from the edge is counted alongside one that -reached the origin. No other source covers both. +CloudFront records the query string for cached and uncached requests. Rainlytics reads the term but +does not include the viewer address in the result. -Who searched stays out of the answer. The delivered field set carries the viewer's address, because -the unique visitor count is hashed from it, and a search count reads the query string alone. Cookies -are left out of the field set altogether. The [log table](../log-table/) page covers the address -column. +## Name the search path and parameter -## Name the search page - -`--path` is what separates a search from every other query string the same log holds. A site's -analytics beacon, its legacy tools and the tracking parameters on inbound links are all in there, -and counting them together answers nothing. - -```bash -rainlytics searches --path /search/ --last 30d -rainlytics searches --path /tools/convert/ --param hanzi --last 30d -``` - -A site with two search boxes gives `--path` once for each, and every row names the box it came from: +The path distinguishes search requests from tracking parameters, beacon events and other query +strings on the site. ```bash -rainlytics searches --path /words/search/ --path /sentences/search/ --last 30d +rainlytics searches --path /search/ --param q --last 30d +rainlytics searches --path /tools/convert/ --param text --last 30d ``` -```text -term section searches redirected --------- ------------------ -------- ---------- -happy /words/search/ 41 38 -happy /sentences/search/ 12 0 -``` - -Two corpora give two answers to the same word. The `section` column names the box a row came from, -written out of the same test that let the row in. Where two of the prefixes overlap (`/guides/` -given alongside `/guides/advanced/`) a row reports the first of them given, and every row is in -exactly one section. - -One `--path` leaves the column out, since every row would carry the same value. One run still reads -both boxes for one query's money, where two runs would be two questions and two bills. +`--param` defaults to `q`. -`--param` names the parameter carrying the term and defaults to `q`. One site can hold several, and -each is its own question. `--redirect-status` names what counts as a search sent to its answer, and -[the section below](#which-statuses-count) covers it. Everything else the rollups take works here -too, `--host`, `--limit` and `--include-bots` among them. See [rollups](../rollups/). +Repeat `--path` for several search pages: -## `redirected` is how many found a page - -A site that answers an exact match by sending the reader straight to its page can read that column -as the searches that found one. The rest produced a list. - -```text -term searches redirected --------- -------- ---------- -好 41 38 -happy 12 0 +```bash +rainlytics searches \ + --path /docs/search/ \ + --path /api/search/ \ + --last 30d ``` -`好` is a word the site publishes, and nearly every search for it goes straight there. Nobody -searching `happy` was sent anywhere, so those twelve readers got a list. - -### Which statuses count +The result then includes a `section` column. If prefixes overlap, the first matching path on the +command line wins. -302, 303 and 307. Those are what a site answers when it sends a reader to the thing they searched -for. +## Redirected searches -301 and 308 are left out, because a permanent redirect is address tidying. A reader gets one -whatever they typed. A site answering `/search?q=happy` with a 308 to `/search/?q=happy` carries the -term on the redirect and again on the request behind it, and counting the 308 reports one reader as -two searches and puts `happy` on the same line as a term the site publishes a page for. A -canonical-host 301 does it again. +`redirected` counts searches that received a temporary redirect. The defaults are 302, 303 and 307. This can represent an exact match that sent the reader directly to a page. -A site whose exact match answers 301 names its own: +Permanent redirects are omitted because canonical URL redirects can count one search twice. Change +the statuses when your search endpoint uses another response: ```bash -rainlytics searches --path /search/ --redirect-status 301,302 --last 30d +rainlytics searches \ + --path /search/ \ + --redirect-status 301,302 \ + --last 30d ``` -One value carrying commas, where `--path` is given again for each path. A path is long and arrives -one at a time, often out of a shell variable. Three status codes are read and typed as one thing. +The access log cannot tell a nonempty result list from an empty one when both return 200. Use a +different status or path if that distinction belongs in analytics. -### An empty result reads as a list +## Stored searches -What the access log cannot tell you is whether that list had anything in it. The record carries the -status, the path and the query, and a ranked list and an empty result are both 200. Separating them -is work for the site being measured, which has to make the two differ in one of those three fields -before any log-based report can pick them apart. A distinct status for an empty result is the -usual lever. +Configure the scheduled question with the same path and parameter: -## The term is decoded once - -CloudFront percent-encodes what it writes and the browser has already encoded the term, so `家` -reaches the log as `%25E5%25AE%25B6`. The rollup reads it with `url_extract_parameter`, which -decodes its own answer, and one further `url_decode` finishes the job. +```typescript +new RollupSummaries(this, "Summaries", { + table, + workgroup, + requests: { + searches: { paths: ["/search/"], param: "q" }, + }, +}); +``` -That is one pass where [pageviews](../rollups/#the-log-is-percent-encoded-twice) needs two, and the -difference is the extract function rather than the data. Two passes here would decode a term -holding a percent sequence twice, and `50%` typed into a search box is the case. +A command that omits `--path` and `--param` adopts the values recorded in the stored summaries. A +different value requires `--query`. -A space arrives as `+` from a form and as `%20` from a hand-written link, and both come back as a -space. The two spellings of one search are one row. +CloudFront encodes the browser's query string again when writing the log. The rollup extracts the +parameter and decodes the term once more. Form `+` and URL `%20` spaces therefore group together. diff --git a/docs/summary-schedule/README.md b/docs/summary-schedule/README.md index 7b752a7..fb825df 100644 --- a/docs/summary-schedule/README.md +++ b/docs/summary-schedule/README.md @@ -1,456 +1,181 @@ # Summary schedule -`RollupSummaries` runs the named questions on a schedule and writes each answer to S3. - -Athena prices per query, and asking the same question twice pays twice. This construct asks each -question once per window and stores the rows, so reading the answer afterwards costs a GET however -many people look. [Reading a precomputed answer](../rollups/#reading-a-precomputed-answer) has what -the command line does with what lands here. +`RollupSummaries` computes common analytics questions on a schedule and writes the answers to S3. ```typescript -import { - CloudFrontLogDelivery, - LogBucket, - LogTable, - QueryWorkgroup, - RollupSummaries, -} from "@kensio/rainlytics/cdk"; - -const logs = new LogBucket(this, "RainlyticsLogs"); -const delivery = new CloudFrontLogDelivery(this, "RainlyticsDelivery", { - distributionId: distribution.distributionId, - logBucket: logs.bucket, -}); -const table = new LogTable(this, "RainlyticsTable", { deliveries: [delivery] }); -const workgroup = new QueryWorkgroup(this, "RainlyticsQueries"); - -new RollupSummaries(this, "RainlyticsSummaries", { table, workgroup }); -``` - -That deploys a bucket, one Lambda function and a schedule for each question on each cadence. The -[summaries](../summaries/) page has the document those schedules produce and where it lands. - -## What runs, and when - -A schedule fires. EventBridge Scheduler invokes the function with the question in its target input. -The function starts the query, waits for it, and puts the rows in the bucket under the key -`summaryKey` builds. Athena does the counting, and the function starts it and stores what came back. - -Two schedules per question by default, one for hours and one for days. Each hourly schedule fires -fifteen minutes into every hour, and each daily one fifteen minutes after midnight UTC. Every window -is UTC, and so is the run that computes it. - -Nothing here is always-on and nothing carries a per-hour floor. Scheduler bills per invocation, -Lambda per millisecond, Athena per byte scanned and S3 per request and per byte. +import { RollupSummaries } from "@kensio/rainlytics/cdk"; -## Why fifteen minutes - -CloudFront delivers a request's log record some time after the request. -[#9](https://github.com/KensioSoftware/rainlytics/issues/9) measured that end to end across 200,074 -records and found a median of 169 seconds and a worst case of 373. An hour's objects have therefore -all landed by four minutes past the next hour, and a run a quarter past has eleven minutes of margin -over the slowest record in that sample. - -A run on the hour would compute every hour before its last records arrived. The tail of each hour -would be missing and every summary would look complete. That is the failure the lag exists to -avoid. - -`lag` moves it. A site that has watched its own delivery and wants fresher answers lowers it. A site -whose logs arrive from several distributions, or whose traffic is bursty, raises it. The lag has to -be a whole number of minutes under an hour, because it decides which minute of the hour a run fires -on and an hour is the shortest window stored. - -```typescript -new RollupSummaries(this, "RainlyticsSummaries", { +const summaries = new RollupSummaries(this, "Summaries", { table, workgroup, - lag: Duration.minutes(25), }); ``` -## Why each run computes two windows +Named CLI commands read these stored answers. Repeated reads use S3 and do not start another Athena +query. -A record CloudFront delivers after its window was computed is invisible to every reader until -something computes that window again. A job that only ever wrote the window that had just closed -never would, and the summary would go on being quietly short for as long as it lived. +## Default schedule -So a run computes the window that has just closed and the one before it, newest first. Both are -written to the keys they were written to before, and each replaces what was there. Recomputing a -window is a re-run of the job. A bug in a rollup is a re-run too. - -`recomputedWindows` moves that count. One computes each window once and never again. Higher numbers -buy more grace at one Athena query each. - -Nothing backfills. A run reaches back as far as `recomputedWindows` and no further, so a window that -closed before the construct was deployed has no summary and reports as `neverComputed` to whatever -reads the bucket. Answering a question about one of those is a `rainlytics` query over raw. - -```typescript -new RollupSummaries(this, "RainlyticsSummaries", { - table, - workgroup, - recomputedWindows: 3, -}); -``` +The construct computes six default questions: -## What it costs +- pageviews +- referrers +- browsers +- status codes +- cache hit ratio +- searches -One Athena query per window per question per run. Athena bills a ten million byte minimum whatever a -query reads. That is $0.00005 a query (the [query](../query/) page has where the per-byte figure -comes from). +Each question runs for hourly and daily UTC windows. A schedule starts 15 minutes after a window +closes and recomputes the two latest closed windows. Recomputing the previous window picks up logs +that CloudFront delivered late. -The six default questions on both cadences, recomputing two windows, make a rollup-only subtotal of -300 queries a day. That is about 45 cents a month. The visitor count on `pageviews` adds 50 queries -and about 8 cents, making the default total 350 queries and about 53 cents. Lambda adds a few cents -and S3 is a rounding error at this scale. The default `recomputedWindows` of 2 doubles the query -cost and leaves the object count alone, since a recomputed window overwrites its own key. +The same construct runs a calendar-report job once a day. It writes closed daily, weekly, monthly +and annual reports. See [Calendar reports](../reports/). -Lowering `recomputedWindows` to 1 halves the Athena bill. Computing hours alone, with -`granularities: ["hourly"]`, is the other lever, at the price of a reader assembling a day out of 24 -objects. +## Create the visitor salt first -A question that counts visitors runs a second query per window. `pageviews` accounts for the 50 -visitor queries in the default total above. [Counting visitors](../visitors/) has what that number -means. +The default pageview question counts visitors and reads `/rainlytics/visitor-salt` from SSM +Parameter Store. Create the `SecureString` before the first scheduled run: -## Reading the query a schedule runs +```bash +aws ssm put-parameter \ + --name /rainlytics/visitor-salt \ + --type SecureString \ + --value "$(openssl rand -hex 32)" +``` -The SQL is written at synthesis by the same builder the `rainlytics` command uses, with the window -left as a placeholder the job fills in when it runs. It is in the CloudFormation template and in the -schedule's target input, so what the job will run can be read without running it. +Run the command in the account and region containing this construct. Pass +`visitorSaltParameter` when you use another name. -A scheduled summary of one hour and a `rainlytics pageviews --last 1h` run over that hour therefore -count it the same way. The [rollups](../rollups/) page has what each question counts. +A delivery without `c-ip` creates summaries without visitor counts and needs no parameter. See +[Counting visitors](../visitors/#run-without-visitor-counts). -The query is fixed at deploy time. A package upgrade that changes what a question counts reaches the -running job when the stack is deployed again, and the same deploy replaces the function's code. +## Give the CLI the bucket name -Each question also takes the narrowing `RollupQueries` takes, per question and by name: +The construct creates a summaries bucket unless you pass one. Output its generated name: ```typescript -const site = { host: "docs.example.com" }; +import { CfnOutput } from "aws-cdk-lib"; -new RollupSummaries(this, "RainlyticsSummaries", { - table, - workgroup, - requests: { - pageviews: site, - searches: { ...site, paths: ["/search/"], param: "term" }, - }, +new CfnOutput(this, "SummaryBucketName", { + value: summaries.bucket.bucketName, }); ``` -The narrowing is recorded in every summary the question produces, so a reader can see that an answer -is a narrower one than they asked for. Only the question's name reaches the key, so two narrowings of -one question want [two rollups with two names](../rollups/#writing-a-rollup-of-your-own). A pair -sharing a name is refused at synthesis, because both would write to one key and whichever ran last -would be the answer. - -## When a run fails - -A query that does not succeed fails the run. The message names the question, the window, Athena's own -reason and the execution id, and it goes to the function's log group. The invocation counts on the -function's `Errors` metric. - -The bytes-scanned cutoff is the failure worth expecting. It is `bytesScannedCutoff` on the -[workgroup](../query-workgroup/), and it is there so one query cannot run up a bill nobody chose. A -scheduled question reads one window and should be nowhere near the ceiling. A run that meets it has -usually lost its partition predicate, which the message says. - -Nobody is watching a scheduled run, and there are two places to look: - -- **The bucket.** A window with no object is one nobody computed. A window with an object holding no - rows saw no traffic. Summaries that stop appearing are the visible half of a job that stopped - working, and the [summaries](../summaries/) page has the three answers a reader meets. -- **The log group.** One of its own, kept for a month by default and moved with `logRetention`. - -CloudFormation names that log group after the stack and the logical id. The name comes out -something like `MyStack-RainlyticsSummariesJobLogs1C6CB09C-8mKvQ2XrTpLd`. The -`/aws/lambda/` a Lambda function's logs usually sit under holds nothing here. The -function's **Monitor** tab in the Lambda console links to the right group, and that is the quickest -way to it. From a terminal, ask the stack: +Set it in the shell that runs Rainlytics: ```bash -aws cloudformation describe-stack-resources --stack-name MyStack \ - --query "StackResources[?ResourceType=='AWS::Logs::LogGroup'].PhysicalResourceId" +export RAINLYTICS_SUMMARY_BUCKET= +rainlytics pageviews --last 7d ``` -A CloudWatch alarm over the error metric is the thing that would tell somebody without their having -to look, and it is the one piece of this that carries a fixed monthly charge. That is why the -construct does not create one. A site that wants the notification more than it wants the constraint -adds an alarm over the function it gets back: +The equivalent flag is `--summaries `. An identity built in CDK can receive read access +with: ```typescript -const summaries = new RollupSummaries(this, "RainlyticsSummaries", { - table, - workgroup, -}); - -summaries.lambda.metricErrors().createAlarm(this, "SummariesFailing", { - threshold: 1, - evaluationPeriods: 1, -}); +summaries.grantReadingSummaries(role); ``` -## Where the summaries go - -A bucket of its own, created by the construct and available as `summaries.bucket`. Its own bucket and -never the log bucket, because the logs expire on a retention measured in months and the answers -computed from them outlive the records. - -It carries no expiry rule. A year of six questions on both cadences is about 55,000 objects of a few -kilobytes, and a summary is the only remaining record of a window once the raw objects have gone. - -Pass `summariesBucket` to write into one of your own. That is worth doing where something outside -this stack reads the answers, such as a static site given read access to one prefix. +That grant adds `s3:GetObject` on stored summaries and KMS decryption when the bucket uses a +customer-managed key. -### Give the command line the generated bucket name +## Configure the questions -Code outside the stack needs the physical bucket name. Keep CloudFormation's generated name and -publish the bucket token as an output from the application stack: +Pass `rollups` to add optional or custom questions: ```typescript -import { CfnOutput } from "aws-cdk-lib"; +import { javascriptErrors, rollups, webVitals } from "@kensio/rainlytics"; -const summaries = new RollupSummaries(this, "RainlyticsSummaries", { +const summaries = new RollupSummaries(this, "Summaries", { table, workgroup, -}); - -new CfnOutput(this, "RainlyticsSummaryBucketName", { - value: summaries.bucket.bucketName, + rollups: [...rollups, javascriptErrors, webVitals], }); ``` -`cdk deploy` prints that value after a successful deployment. Its `--outputs-file` option writes -the stack outputs as JSON as well: +Optional browser questions stay out of the defaults because an access-log-only site would pay for +empty Athena queries. -```bash -cdk deploy AnalyticsStack --outputs-file cdk-outputs.json - -export RAINLYTICS_SUMMARY_BUCKET="$( - jq -r '.AnalyticsStack.RainlyticsSummaryBucketName' cdk-outputs.json -)" -``` - -Put `outputsFile` in the application's `cdk.json` to refresh the file on every deploy without -repeating the option. CDK groups a multi-stack deployment by stack name, and the `jq` path above -selects one deployment explicitly. The Rainlytics command then keeps its existing permissions. It -reads summary objects from S3 and makes no CloudFormation lookup. - -`summariesBucketName` gives a new deployment a fixed physical name where that tradeoff is useful. -Changing it later replaces the bucket. The old bucket is retained by default and keeps the stored -history. A fixed name can also block a redeploy after a retained bucket survives rollback. The -[log bucket](../log-bucket/#when-a-deploy-rolls-back) has the same retention and naming tradeoff in -more detail. - -## Props - -| Prop | Default | What it decides | -| ---------------------- | -------------------------- | -------------------------------------------------- | -| `table` | required | The Glue table the questions read. | -| `workgroup` | required | Where the queries run, and their cutoff. | -| `rollups` | the six default questions | What to compute. | -| `requests` | none | What each question covers. | -| `granularities` | `["hourly", "daily"]` | Which windows to compute. | -| `lag` | 15 minutes | How long after a window closes a run fires. | -| `recomputedWindows` | 2 | How many closed windows a run computes. | -| `summariesBucket` | one is created | Where the answers land. | -| `summariesBucketName` | CloudFormation-generated | The created bucket's physical name. | -| `visitorSaltParameter` | `/rainlytics/visitor-salt` | The SSM parameter holding the visitor salt secret. | -| `timeout` | 5 minutes | How long one run may take. | -| `logRetention` | a month | How long the function's logs are kept. | -| `schedulePrefix` | `rainlytics-` | What each schedule's name begins with. | - -A default deployment reads the salt parameter. `pageviews` counts visitors, and it is one of the six -questions above. The `SecureString` has to be in Parameter Store before the first run ([creating the -secret](../visitors/#creating-the-secret) has the command). A deployment that wants none passes -`rollups` without a question that counts visitors, and never reads the parameter. - -## Two deployments in one account - -A schedule's name is unique within its group, and every schedule here goes in the account's default -group. A second Rainlytics deployment in the same account and Region therefore meets the first one's -`rainlytics-pageviews-hourly` and fails at deploy time. `schedulePrefix` is how the second one says -which it is: +Use `requests` to store a narrowed version of a question: ```typescript -new RollupSummaries(this, "RainlyticsSummaries", { +const summaries = new RollupSummaries(this, "Summaries", { table, workgroup, - schedulePrefix: "docs-", + requests: { + searches: { paths: ["/search/"], param: "q" }, + }, }); ``` -The same holds for `workgroupName` on the [workgroup](../query-workgroup/) and `databaseName` on the -[table](../log-table/). One deployment per account reads well by default, and a second one names -itself. - -## Questions of your own +The summary records this narrowing. A CLI command that omits the same filters adopts the stored +configuration. A command that requests different filters stops and suggests `--query`. -A rollup a site wrote is scheduled like a shipped one, and its SQL comes from the same builder: +## Configure windows and reports ```typescript -new RollupSummaries(this, "RainlyticsSummaries", { +import { Duration } from "aws-cdk-lib"; + +const summaries = new RollupSummaries(this, "Summaries", { table, workgroup, - rollups: [...rollups, countries], + granularities: ["hourly", "daily"], + lag: Duration.minutes(30), + recomputedWindows: 3, + reportTimeZone: "Europe/London", + reportWeekStartsOn: "monday", }); ``` -[Writing a rollup of your own](../rollups/#writing-a-rollup-of-your-own) has what a question has to -do. The one rule this construct adds is that the question builds its `WHERE` clause with `rowsFor`, -because that is what writes the window the job fills in. A query that reaches Athena without one is -refused before it is sent, since Athena would take it and read every partition the table projects. +Increase `lag` if logs regularly arrive after the scheduled run. Increase `recomputedWindows` when +late delivery extends further back. Both changes increase the number of queries or delay fresh +answers. -## Permissions for a scoped deploy role +## Cost -Skippable on an account whose CloudFormation execution role holds `AdministratorAccess`. +Every schedule invocation starts a Lambda function and at least one Athena query. S3 stores the +small JSON result. Scheduler, Lambda, Athena and S3 are all usage-priced. -This construct creates more kinds of resource than the rest of the pipeline. A deploy of the -defaults writes ten schedules, one function, one log group, two roles with an inline policy each, -and a bucket with a bucket policy. `scheduler:` is the prefix to check first. A role that has -deployed anything else in the account usually holds the rest already. +The default six questions, two granularities and two-window recomputation run 300 rollup queries a +day. The pageview visitor count adds 50 more. Athena bills each query at its minimum even when the +window has no rows. At the standard 10 MB minimum and $5 per TB, those 350 queries are about 53 +cents in an average month before traffic pushes a query above the minimum. -```typescript -import { PolicyStatement } from "aws-cdk-lib/aws-iam"; - -new PolicyStatement({ - sid: "TheRainlyticsSummarySchedules", - actions: [ - "scheduler:CreateSchedule", - "scheduler:GetSchedule", - "scheduler:UpdateSchedule", - "scheduler:DeleteSchedule", - ], - resources: [ - `arn:aws:scheduler:${region}:${account}:schedule/default/${schedulePrefix}*`, - ], -}); -``` +Optional questions add 50 queries a day under the same defaults. Calendar reports add period-wide +queries where stored summaries cannot be combined correctly. -`default` is the schedule group every schedule here goes in, and the names after it begin with -`schedulePrefix` (`rainlytics-` unless it was passed). The wildcard therefore covers one -deployment's schedules and leaves whatever else the account has scheduled alone. +There is no always-on resource or reserved capacity. A deployment with schedules still has the +minimum-query cost even when its site receives no traffic. -It matches on the prefix and nothing else. A second deployment that called itself `rainlytics-docs-` -would sit inside a statement quoting `rainlytics-*`, and the first deployment's role could update -and delete its schedules. Give the second one a prefix that is not the opening of the first one's. +## Detect failed runs -The function, its log group and the two roles: +A failed query stops that Lambda invocation. Other questions continue because each question has its +own schedule. Check the summary function's CloudWatch log group and Lambda error metric when a +stored window is missing. -```typescript -new PolicyStatement({ - sid: "TheRainlyticsSummaryFunction", - actions: [ - "lambda:CreateFunction", - "lambda:GetFunction", - "lambda:UpdateFunctionCode", - "lambda:UpdateFunctionConfiguration", - "lambda:DeleteFunction", - ], - resources: [`arn:aws:lambda:${region}:${account}:function:${stackName}-*`], -}); +Rainlytics does not create an alarm by default. CloudWatch alarms can add a fixed monthly charge, +so the deployment decides whether to add one. -new PolicyStatement({ - sid: "TheRainlyticsSummaryLogs", - actions: [ - "logs:CreateLogGroup", - "logs:PutRetentionPolicy", - "logs:DeleteLogGroup", - ], - resources: [`arn:aws:logs:${region}:${account}:log-group:${stackName}-*`], -}); +## Run more than one deployment -new PolicyStatement({ - sid: "TheRainlyticsSummaryLogGroups", - actions: ["logs:DescribeLogGroups"], - resources: ["*"], -}); +Schedule names are unique in an account and region. Give another deployment a distinct prefix: -new PolicyStatement({ - sid: "TheRainlyticsSummaryRoles", - actions: [ - "iam:CreateRole", - "iam:GetRole", - "iam:DeleteRole", - "iam:PutRolePolicy", - "iam:GetRolePolicy", - "iam:DeleteRolePolicy", - "iam:AttachRolePolicy", - "iam:DetachRolePolicy", - "iam:PassRole", - ], - resources: [`arn:aws:iam::${account}:role/${stackName}-*`], +```typescript +const summaries = new RollupSummaries(this, "DocsSummaries", { + table, + workgroup, + schedulePrefix: "docs-rainlytics-", }); ``` -`logs:DescribeLogGroups` is the one that has to be `*`. CloudFormation reads a log group back with -it, and IAM gives the action no resource type at all, so a statement naming a log group authorises -it for no request. The other three actions are scoped to the group. - -`iam:PassRole` is the one that looks like surplus. `lambda:CreateFunction` hands Lambda the role the -function runs as, and `scheduler:CreateSchedule` hands Scheduler the role it invokes through. A -policy that creates both roles and stops there fails on the first of those two calls. -`iam:AttachRolePolicy` is for `AWSLambdaBasicExecutionRole`, the managed policy CDK attaches to -every function's role. - -The function, its log group and both roles are left unnamed, and CloudFormation names each of them -after the stack and the logical id. That is where `${stackName}-*` comes from above, and the -truncation trap on the [log bucket](../log-bucket/) page applies to all three. Check the prefix -against the names a deploy created rather than against the stack name alone. - -The created bucket takes the S3 permissions on that same page against its own ARN, its bucket policy -included (`enforceSSL` writes one). A deployment passing `summariesBucket` creates no bucket and -needs none of them. - -The function's code goes up under a different role. `cdk deploy` uploads the asset with the -bootstrap file publishing role before CloudFormation runs. A deploy that fails on the upload is a -bootstrap question. - -Read the whole list as inferred from what the construct creates. The resource counts above come -from synthesising the defaults, and no deploy has run under a role narrower than -`AdministratorAccess`. So the actions have never been tested against the failure a missing one would -cause. - -## Permissions for a scheduled run - -The deploy role has no part in these. The construct grants them itself, onto the role the function -runs as, and a reader under a narrowed deploy role has no policy to write for them. They are here -because a service control policy denies an unlisted prefix whichever role sent the call, and the -deploy list above is half of what such an account has to allow. - -One run sends: - -- `athena:StartQueryExecution`, `StopQueryExecution`, `GetQueryExecution` and `GetQueryResults` on - the workgroup, with `athena:GetWorkGroup` alongside them (Athena reads the workgroup's own - configuration on the way to running a query in it). -- `glue:GetDatabase`, `glue:GetTable` and `glue:GetPartitions` on the catalog, the database and the - table. -- `s3:GetObject`, `s3:GetBucketLocation` and `s3:ListBucket` on the log bucket and its objects. - Athena lists the prefixes a partition predicate selected before it reads anything in them. -- On the workgroup's results bucket and its objects, what CDK's `grantReadWrite` writes. - `s3:GetObject*`, `s3:GetBucket*`, `s3:List*`, `s3:DeleteObject*`, `s3:PutObject` with its - `LegalHold`, `Retention`, `Tagging` and `VersionTagging` variants, and `s3:Abort*`. Athena writes - each query's output there as the caller and reads it back to answer `GetQueryResults`. -- The same `PutObject` family and `s3:Abort*` again on the summaries bucket's objects, from - `grantPut`. -- `ssm:GetParameter` on the visitor salt parameter. - -Those are the statements on the deployed role read back off a synthesised template, rather than a -list inferred from the code. - -An account that allows the deploy prefixes and stops there deploys cleanly and fails on the first -schedule that fires. Nobody is watching that run. The message lands in the function's log group and -the summaries never start appearing. That is the harder of the two failures to attribute. - -`ssm:GetParameter` reads a parameter no template creates. The salt is a `SecureString`, and -CloudFormation writes `String` and `StringList` parameters only. Somebody puts it there by hand -before the first run that counts visitors, and [creating the -secret](../visitors/#creating-the-secret) has the command. +Also give its `LogTable` database and `QueryWorkgroup` unique names. Avoid prefixes where one is the +start of another if IAM policies scope schedule access by prefix. diff --git a/docs/visitors/README.md b/docs/visitors/README.md index 4c77864..cdb7585 100644 --- a/docs/visitors/README.md +++ b/docs/visitors/README.md @@ -1,220 +1,140 @@ # Counting visitors -A visitor is one browser on one day. Rainlytics counts them from the viewer address CloudFront -records, hashed under a salt that changes every day. +Rainlytics defines a visitor as one browser identity inside one reporting period. It derives the +identity from the viewer address and user agent in the CloudFront access log. ```json "visitors": { "distinct": 317, "additive": false } ``` -That field rides on a [rollup summary](../summaries/) alongside the rows, on the questions that -count pages. `additive: false` is there because two days of visitor counts do not add up, and the -rest of this page is why. +## What the number means -## What the number is +The count describes browser connections, not known people: -The count is over the pageviews the summary's question covers. A summary narrowed to `/blog/` -reports the visitors to `/blog/`, and one narrowed to a host reports that host's. Views and visitors -are two numbers over one set of rows. +- two devices usually count twice +- identical browsers behind one household address can count once +- carrier-grade NAT can merge many people +- changing VPN or network can split one person +- records without a viewer address cannot count a visitor -Automated traffic is left out, the way it is everywhere else. Bots are most of a quiet site's -traffic and each crawler would otherwise arrive as a visitor a day. +The number is most useful when compared with the same site and configuration over time. -## What it stands for +Automated traffic is omitted by default. The user-agent filter cannot identify every bot, and a +client can send any user agent it chooses. -One browser on one day, in one place. +## How the identifier is built -- **Two devices are two visitors.** A phone on the train and a laptop at a desk carry two addresses. -- **A household is one visitor**, where two people behind one router run the same browser. The user - agent goes into the hash for this reason and separates a phone from a laptop on the same address, - and it cannot separate two identical iPhones. -- **A mobile carrier is fewer visitors than it should be.** Carrier-grade NAT puts thousands of - people behind one address, and the user agent is what splits them, imperfectly. -- **A VPN moves somebody**, and the same person on and off one is two visitors. -- **A record with no address is nobody.** CloudFront started recording addresses for Rainlytics in - [#73](https://github.com/KensioSoftware/rainlytics/issues/73). Every day before that delivery - change counts zero visitors. - -So it is a measure of browsers rather than of people, and the number moves with how a site's -readers connect. Every privacy-preserving analytics product reports the same measure, and each of -them counts a slightly different set of browsers. Compare the number against itself over time. - -## What cannot be added up - -The salt changes at midnight UTC. The same browser carries one identifier today and a different one -tomorrow. - -A day of them counts. Two days added together count everybody who came back twice over, and a month -of them is a figure about nothing. Thirty summaries each carrying `"distinct": 429` are thirty -numbers `jq` will happily sum, and the total describes nobody. - -That is what `additive: false` says in the document, and what the `VisitorCount` wrapper says to -TypeScript. A month of visitors is a query over raw, under the salt that month was counted with. - -Hours work differently. Every hour of a day shares that day's salt, and the hourly summaries of a -day count the same identifiers the daily summary counts. They still fail to add, because somebody -who came back after lunch appears in two of them. The daily summary is the answer for a day. - -## The identifier +Athena hashes the period salt, viewer address and user agent: ```sql -to_hex(sha256(to_utf8(concat(, '|', c_ip, '|', cs_user_agent)))) +to_hex(sha256(to_utf8(concat(, '|', c_ip, '|', cs_user_agent)))) ``` -Athena computes it while counting, and every digest dies with the query that made it. The summary -holds the count and no more. - -The three parts are joined by `|`, which an address cannot contain. The text hashed for one address -and user agent therefore belongs to that pair alone. - -SHA-256 rather than the faster `xxhash64`. Both are in Athena engine version 3, and a 64-bit -non-cryptographic digest is forgeable by anybody holding one. At the volumes a site of this size -produces, the speed makes no difference worth having. +The digest exists inside the query. Summaries store the final count only. -## Where the salt lives - -The salt for a day is derived from one secret and the date: +For daily summaries, the salt comes from an HMAC of one deployment secret and the UTC date: ```text -salt(day) = HMAC-SHA256(secret, "rainlytics/visitor-salt/1/" + day) +HMAC-SHA256(secret, "rainlytics/visitor-salt/1/" + date) ``` -The secret is a `SecureString` in SSM Parameter Store. The Lambda that computes the summaries reads -it once per run and derives the salt for each window it is computing. Four things follow, and they -are the four the decision in -[#53](https://github.com/KensioSoftware/rainlytics/issues/53#issuecomment-3576795104) asked for. - -**Every record of a day counts under one salt.** The day comes from the window, and every window -inside a day gives the same date. +The same date produces the same salt during recomputation. Another date produces a different salt. +Calendar reports derive a separate salt for their full period. -**Tomorrow is a different salt.** The date is in the message the HMAC is taken over. +The salt used by Athena appears as a literal in Athena query history. The deployment secret does +not. A period salt cannot derive the secret or another period's salt. -**A re-run of a day reproduces it.** The date comes from the window being computed and never from -the clock. A window recomputed next week therefore writes the count that was there before it. The -[summary schedule](../summary-schedule/) recomputes a trailing window on every run for exactly this, -and a salt taken from the clock would make the second run disagree with the first. +## Create the secret -**A reader of the log bucket cannot get it.** The secret lives in Parameter Store alone, away from -the bucket, the summaries, the CloudFormation template and the schedule that carries the query. It -is encrypted at rest under the `aws/ssm` managed key, and reading it takes `ssm:GetParameter` on -that one parameter. - -The secret is meant to stand rather than rotate. Replacing it makes every day from then on count -somebody new, and makes every day before it uncountable. The date is what rotates. - -### The salt reaches Athena as text - -Athena takes no secret of its own. The salt goes into the statement as a quoted literal, and a copy -of it then lives wherever Athena keeps its -[query history](https://docs.aws.amazon.com/athena/latest/ug/querying-keeping-query-history.html) -(45 days, behind `athena:GetQueryExecution` on the workgroup). CloudTrail -[records the query string as `***OMITTED***`](https://docs.aws.amazon.com/athena/latest/ug/monitor-with-cloudtrail.html) -for `StartQueryExecution`, and the statement reaches S3 nowhere. - -This is why the statement carries a day's salt and never the secret. HMAC is built so that a key -cannot be recovered from a message and its digest. A salt read out of query history therefore covers -the days it appears for, and says nothing about any other day or about the secret. - -## Creating the secret - -Nothing creates it for you. CloudFormation writes `String` and `StringList` parameters and no -`SecureString`, and a construct that generated a secret at synthesis would put it in a template, -which is the one place it must not be. +Create one SSM Parameter Store `SecureString` in the account and region containing the summary +jobs: ```bash -aws ssm put-parameter --name /rainlytics/visitor-salt --type SecureString --value "$(openssl rand -hex 32)" +aws ssm put-parameter \ + --name /rainlytics/visitor-salt \ + --type SecureString \ + --value "$(openssl rand -hex 32)" ``` -Run it once per deployment, in the account and region the summaries run in. The -[summary schedule](../summary-schedule/) construct takes `visitorSaltParameter` for a name of your -own, and grants the job `ssm:GetParameter` on whichever one it was given. +CloudFormation cannot create a `SecureString`, and generating the value during synthesis would put +the secret in the template. + +Keep this secret for the lifetime of the deployment. Replacing it breaks continuity and prevents +past periods from being recomputed with their original identity set. Pass another parameter name +with `visitorSaltParameter` on `RollupSummaries`. + +## Combining visitor counts -A run that meets no parameter fails and says so, naming the parameter and printing that command. -[Which questions carry a count](#which-questions-carry-a-count) says which runs read it. +Two hourly counts can contain the same browser. Two daily counts deliberately use different salts. +Adding either pair double-counts returning visitors. -## Which questions carry a count +The summary marks this rule with `additive: false`. The CLI refuses to add visitor values across +stored windows. Use a calendar report or `--query` for one identity set over a larger period. -`pageviews` alone, and it is one of the six questions a deployment gets when it passes no `rollups` -of its own. A default deployment therefore reads the salt parameter, and the secret has to be there -before its first run. [Running without a visitor count](#running-without-a-visitor-count) has the -deployment that reads no parameter at all. +## Questions that count visitors -A rollup says it counts with `countsVisitors`: +The default `pageviews` rollup counts visitors. A custom rollup opts in with +`countsVisitors: true`: ```typescript import { pageviews, type Rollup } from "@kensio/rainlytics"; -const blogVisitors: Rollup = { +const articlePageviews: Rollup = { ...pageviews, - name: "blog-pageviews", + name: "article-pageviews", countsVisitors: true, }; ``` -The count is always over pages, whatever the question beside it counts. A summary of status codes -carrying one would report a number about rows it never looked at, and the field is left out of every -question that counts something else. Absent and `{ "distinct": 0 }` mean different things, and a -reader can tell them apart. +The visitor count always covers pageview rows under the same host and path filters. A question that +counts another kind of event should normally omit it. -It costs one extra Athena query per window per run. The six default rollup queries on both cadences, -recomputing two windows, make a subtotal of 300 queries a day and about 45 cents a month. The visitor -count on `pageviews` adds 50 queries and about 8 cents, making the default total 350 queries and -about 53 cents. The [summary schedule](../summary-schedule/#what-it-costs) page has the arithmetic. +Counting visitors adds one Athena query per scheduled window. Under the default two granularities +and two-window recomputation, this is 50 queries a day. -## Running without a visitor count +## Run without visitor counts -A site that delivers no viewer address counts no visitors, and nothing else about it changes. +Omit the viewer address from log delivery: ```typescript import { logFieldNamesWithoutAddress } from "@kensio/rainlytics"; -new CloudFrontLogDelivery(this, "RainlyticsDelivery", { +const delivery = new CloudFrontLogDelivery(this, "Delivery", { distributionId: "E1EXAMPLE1234", logBucket: logs.bucket, fields: logFieldNamesWithoutAddress, }); ``` -That is the only line a site changes. The [log table](../log-table/) describes what the delivery -writes and the [summary schedule](../summary-schedule/) reads the table. Both follow. The schedule -computes the same six questions with the count off, needs no salt parameter, and is granted no -`ssm:GetParameter`. Summaries carry no `visitors` field, which a reader tells apart from a count of -zero. +`LogTable` then has no `c_ip` column. `RollupSummaries` computes the default questions without a +visitor count, needs no salt parameter and receives no `ssm:GetParameter` permission. -A deployment naming its own questions says so per question: +For an explicit question list, remove visitor counting from a rollup: ```typescript import { pageviews, referrers, withoutVisitorCount } from "@kensio/rainlytics"; -new RollupSummaries(this, "RainlyticsSummaries", { +new RollupSummaries(this, "Summaries", { table, workgroup, rollups: [withoutVisitorCount(pageviews), referrers], }); ``` -A question that counts visitors over a table with no address is refused at synthesis, naming the -question. Left alone it would run once an hour against a column the table has never heard of. +A question that requires visitor addresses is rejected during synthesis when the table has no +`c_ip` field. -The choice sits on the delivery because the delivery is what writes the raw store. Turning the -address off later leaves every address already written where it is, until the [log -bucket](../log-bucket/) expiry reaches it. +## Raw addresses -## Where the addresses are +The default log bucket stores viewer addresses in clear text for its retention period. The salt +protects derived identifiers, not the source rows. Anyone who can read the log bucket can read the +addresses. -The raw log bucket holds viewer addresses in the clear, for as long as it holds anything. That is -the price [#53](https://github.com/KensioSoftware/rainlytics/issues/53) paid for a visitor count, -and the [log bucket](../log-bucket/) page has the expiry that decides how long it lasts. A site -running the field set above has none of them to keep. - -The salt protects the identifier and never the source. Anyone who can read the log bucket has the -addresses themselves, at better resolution than any digest would give them. +Changing the field set affects new log objects only. Existing addresses remain until the log bucket +lifecycle expires them. diff --git a/docs/web-vitals/README.md b/docs/web-vitals/README.md index 2c7cd54..8a17048 100644 --- a/docs/web-vitals/README.md +++ b/docs/web-vitals/README.md @@ -1,12 +1,12 @@ # Web Vitals -`web-vitals` reports the 75th percentile of each Web Vital collected by the browser beacon. +The `web-vitals` command reports the 75th percentile of each performance measurement collected by +the browser module. ```typescript import { rollups, webVitals } from "@kensio/rainlytics"; -import { RollupSummaries } from "@kensio/rainlytics/cdk"; -new RollupSummaries(this, "RainlyticsSummaries", { +new RollupSummaries(this, "Summaries", { table, workgroup, rollups: [...rollups, webVitals], @@ -26,86 +26,47 @@ lcp 2180 41 ttfb 164 47 ``` -LCP, FCP and TTFB are milliseconds. CLS is a unitless score. `samples` is the number of numeric -measurements behind each percentile. +LCP, FCP and TTFB use milliseconds. CLS is a unitless score. `samples` shows how many numeric +measurements produced the percentile. -## A site opts into it +## Enable collection -The six default rollups work from access-log fields every deployment has. This one reads rows from -the optional browser beacon. A site without those rows would pay for an empty Athena query on every -window, so `webVitals` stays outside the exported `rollups` list. - -Adding it on both cadences with the default two-window recomputation makes 50 more Athena queries a -day. At Athena's ten million byte minimum, that comes to about 8 cents a month. The Lambda and S3 -charges add a few cents or less at this scale. - -The `web-vitals` command ships whether a deployment computes the rollup or not. Without its summary, -the command reports that the window was never computed. `--query` runs the same question from raw -logs. +```typescript +import { startBeacon } from "@kensio/rainlytics/beacon"; +import { reportVitals } from "@kensio/rainlytics/beacon/vitals"; -## One percentile per vital +const beacon = startBeacon(); +reportVitals(beacon); +``` -The rollup calculates p75 with Athena's `approx_percentile`. Web Vitals thresholds use p75 because -it describes the experience of most visits without letting a small number of outliers decide the -result. One percentile keeps the answer aligned with those thresholds, without a p50 or p95 spread. +The optional rollup stays outside the defaults because a site without vital events would pay for +empty Athena queries. Adding it under the default schedule adds 50 queries a day. -The answer has one row for each of `lcp`, `cls`, `fcp` and `ttfb`. Route changes, JavaScript errors, -custom event names, negative values and non-numeric values are filtered out before the percentile -is calculated. +## Calculation -INP is absent. Rainlytics leaves its collection to the `web-vitals` library, and a site reporting -`inp` through that integration still needs a query of its own for now. The [browser -beacon](../beacon/) page covers the collection decision. +The rollup uses Athena `approx_percentile` to calculate p75 separately for TTFB, FCP, LCP and CLS. +It ignores route events, errors, custom event names, negative values and invalid numbers. -`approx_percentile` is approximate. Athena engine changes can also move its answer slightly. Use it -for the performance classification it was designed for, with the sample count beside it. +The answer is site-wide. Use `--host` when one distribution serves several hostnames. `--path` +selects the beacon collection path, not the page reported inside each event. -## The answer is site-wide +INP is absent from the shipped rollup. A site can collect it through `web-vitals` and define a +custom rollup. -Each row groups every reported page together. A quiet site already has few samples in an hour, and -grouping again by page would leave many percentiles resting on one or two visits. `--host` separates -sites where one distribution serves several hostnames. +## Read a useful sample -`--path` has a different job here. It names the collection path in the access log. The default is -`/_rainlytics`, matching `BeaconPath`. A deployment using another collection path records that -narrowing on its summaries: +A percentile from one sample is that sample. Small hourly samples can move sharply, so read +`samples` beside `p75`. -```typescript -new RollupSummaries(this, "RainlyticsSummaries", { - table, - workgroup, - rollups: [...rollups, webVitals], - requests: { "web-vitals": { paths: ["/_measure"] } }, -}); -``` - -The matching query is: - -```bash -rainlytics web-vitals --path /_measure --last 7d --query -``` - -## Small windows move quickly - -A percentile from one sample is that sample. With only a handful, each visit can move p75 a long -way. `samples` makes that visible without choosing a universal minimum that would hide a quiet -site's data. - -Hourly summaries suit spotting a sudden regression, but the count belongs beside the number. A -longer `--query` run gives a steadier percentile when the hourly sample is sparse: +Run one longer Athena query for a steadier value: ```bash rainlytics web-vitals --last 30d --query ``` -## Combining stored windows - -The p75 of two hours cannot be recovered from the two p75 values. The raw measurements and their -distribution have already been reduced. Averaging the values or weighting them by `samples` would -produce a different statistic from p75 for the combined visits. - -For that reason, `webVitals` declares no `totals`. The command reads one stored window and refuses a -span covering several. Use `--query` for one percentile over the whole span. +Several stored p75 values cannot be combined into the p75 for their combined raw measurements. +Weighting or averaging them produces another statistic. The command therefore reads a single stored +window or requires `--query` for a range covering several windows.