Skip to content

tsm2arc

CI Release Go Reference License

Migrate InfluxDB 1.x (1.7/1.8) and 2.x (2.0–2.7) data into Arc by reading TSM and WAL files directly off disk — no running influxd required. InfluxDB 3 (Parquet engine) sources are supported experimentally — see InfluxDB 3 sources.

The on-disk TSM/WAL format is the same across 1.x and 2.x; tsm2arc auto-detects the layout. For 2.x it resolves bucket IDs to readable names from influxd.bolt and skips InfluxDB's internal system buckets (e.g. _monitoring, _tasks).

Built for the case where InfluxDB data sits on cold/unmounted volumes (e.g. EBS snapshots) that can be mounted read-only but are not served by any InfluxDB instance. tsm2arc decodes the TSM/WAL block codecs natively in Go, reconstructs multi-field line-protocol points, and streams them into Arc's /api/v1/import/lp endpoint in resumable, size-bounded, parallel chunks.

Reusable beyond Arc. The extraction side produces standard InfluxDB line protocol; the Arc sink is just the first sink. The TSM/WAL decoder is Apache-2.0 and can be reused to migrate InfluxDB 1.x data into other systems — contributions of new sinks (ClickHouse, QuestDB, TimescaleDB, …) are welcome. See CONTRIBUTING.md.

Install

Download a prebuilt binary from Releases (Linux/macOS/Windows, amd64/arm64), or:

# from source (Go 1.25+)
go install github.com/basekick-labs/tsm2arc/cmd/tsm2arc@latest

# container
docker run --rm ghcr.io/basekick-labs/tsm2arc:latest --version

Each release ships an SBOM (SPDX) and checksums.txt; the container image is multi-arch (linux amd64/arm64) on GHCR.

Status

Feature-complete. Capabilities:

Capability State
Native TSM reader + field-rejoin + LP encode + --dry-run ✅
Chunked gzip POST to Arc /api/v1/import/lp (per-DB routing) ✅
SQLite checkpoint + crash-safe resume ✅
WAL (.wal) reader — merged with TSM per shard ✅
Parallel workers (--workers) + live progress reporting ✅
InfluxDB 2.x layout auto-detection + bucket-name resolution ✅
Measurement rename map + invalid-name policy (fail/skip/map) with a checkpoint audit trail ✅
InfluxDB 3.x Parquet-engine object stores (local disk) — see InfluxDB 3 sources 🧪 experimental

The TSM/WAL codecs (timestamp, float, integer, unsigned, boolean, string) are validated against the real InfluxDB 1.7.11 encoder in unit tests (internal/tsm/decode_test.go, file_test.go, internal/wal/wal_test.go) and cross-checked against real InfluxDB 2.7 data.

Quick start (dry-run)

go build ./cmd/tsm2arc

# InfluxDB 1.x — point at the data dir (or its parent; layout is auto-detected)
./tsm2arc --datadir /var/lib/influxdb --dry-run --sample 10

# InfluxDB 2.x — point at the v2 root (~/.influxdbv2 or the mounted volume);
# engine/data, engine/wal, and influxd.bolt (for bucket names) are auto-detected
./tsm2arc --datadir /var/lib/influxdb2 --dry-run --sample 10

Dry-run discovers shards, decodes every block, reconstructs points, and prints per-database/bucket counts + sample line protocol — without writing to Arc. This is the safe first contact with the source data.

tsm2arc auto-detects 1.x vs 2.x. You can point --datadir at the version's natural root and it resolves the rest:

  • 1.x: the InfluxDB root (containing data/), or …/data directly.
  • 2.x: the v2 root (containing engine/ and influxd.bolt), …/engine, or …/engine/data directly. Bucket names come from influxd.bolt; override its location with --bolt if it lives elsewhere. Without it, buckets fall back to their 16-hex IDs as database names.

Load into Arc

./tsm2arc \
  --datadir /mnt/influxdb/data \
  --waldir  /mnt/influxdb/wal \     # IMPORTANT: include the WAL (see below)
  --arc-url https://arc.example.net \
  --token   "$ARC_TOKEN" \          # admin-tier token (or ARC_TOKEN env)
  --verbose

# options:
#   --waldir DIR             InfluxDB WAL directory — read un-flushed data too
#                            (auto-detected for 2.x as engine/wal)
#   --bolt PATH              InfluxDB 2.x influxd.bolt for bucket names (auto-detected)
#   --workers N              concurrent shards to migrate (default 2)
#   --db-map old=new         rename a source DB/bucket to a different Arc DB (repeatable)
#   --database-filter db     migrate only this source DB/bucket (repeatable)
#   --measurement-map old=new       rename a source measurement (repeatable; new must
#                                   satisfy Arc's name rule — see below)
#   --measurement-map-file PATH     file of measurement renames, one old=new per line
#                                   (blank lines and # comments ignored)
#   --on-invalid-measurement MODE   what to do with a measurement name Arc would
#                                   reject, after --measurement-map is applied:
#                                   fail (default) | skip (drop + report) | map
#                                   (deterministic auto-rename)
#   --chunk-bytes SIZE       raw-LP per import request; bytes or a suffix like
#                            450MB (must be <500MB; default 450MB)
#   --checkpoint PATH        SQLite resume store (default tsm2arc.checkpoint.db)
#   --start / --end          RFC3339 UTC time filters
#   --precision ns|us|ms|s   precision value sent to Arc (default ns; tsm2arc always emits ns)
#   --include-internal       also migrate InfluxDB 1.x's _internal database
#                            (2.x system buckets _monitoring/_tasks are always skipped)
#   --inflight N             concurrent import POSTs per shard (default 1);
#                            commits stay strictly ordered, so resume is exact
#                            at any value. >1 widens the tagless crash-duplicate
#                            bound to <=N chunks/shard and multiplies host and
#                            Arc memory (see the runbook)
#   --pipeline               overlap extraction with upload (default true; =false
#                            reverts to serial send and saves one chunk buffer
#                            of memory per worker)
#   --shard-split N          max concurrent merge tasks per shard (default 1).
#                            Output is byte-identical to serial, so resume works
#                            across different values. Requires --merge-memory
#   --merge-memory SIZE      per-shard admission budget for concurrent merges
#                            (e.g. 24GB). Concurrent merges are bounded by
#                            memory, not cores: a merge holds ~one decoded block
#                            per (file x field) stream. Oversized tasks run alone
#   --index-cache SIZE       per-shard budget for cached TSM file indexes
#                            (default 2GB; 0 disables). Avoids re-parsing every
#                            file's index once per series during extraction —
#                            a large CPU cost on shards with many series.
#   --analyze                index-only shard profile (fast; nothing decoded or
#                            sent): series/file/key counts per shard and window-
#                            split profiles for the largest merge runs
#   --analyze-runs N         text output: show the N largest runs per shard
#                            (default 5, labeled when it truncates; 0 = all)
#   --format text|json       with --analyze: json emits EVERY run of every shard
#                            as one JSON document on a pure stdout (for tooling)
#   --send-timeout DUR       per-attempt deadline for one import POST (default
#                            auto: 2m + 1s per MiB of --chunk-bytes). Bounds dead
#                            connections; too low re-sends healthy imports
#   --stall-warn DUR         warn when a shard makes no progress for this long
#                            (default auto: send-timeout + 2m, min 5m)
#   --redact                 with --analyze: replace database, retention policy,
#                            and series names with stable hashed identifiers so
#                            the report can be shared outside your organization
#   --dry-run                extract + count, do not write to Arc
#   --sample N               print N sample LP lines per DB in --dry-run (default 5)
#   --verbose                per-shard / per-chunk logging
#   --version                print version and exit

Parallelism and the --workers knob

--workers N migrates N shards concurrently. Shards are fully independent (each has its own chunk sequence and checkpoint rows), so this scales cleanly. A live heartbeat reports shards-done, chunks, rows, MB, and throughput as the migration runs.

Size --workers against the Arc node's memory, not the migration host's: Arc's import endpoint buffers each request fully in memory (~chunk-bytes decompressed + parsed records), so peak transient Arc-side memory is roughly workers × ~1–1.3 GB at the default 450 MB chunk size. The default of 2 is conservative; raise it deliberately if the Arc node has headroom (the customer's big dedicated migration host is rarely the bottleneck — Arc is). Resume and correctness are unaffected by the worker count.

Memory profile

tsm2arc extracts by streaming one series at a time and, within a series, one TSM block at a time. It indexes only the key list up front, then merges each series straight off lazy block cursors. So on the migration host:

  • Extraction / --dry-run is near-constant memory — it does not scale with the shard, the dataset, or the largest series. Measured on one series written as 1000-value blocks:

    values in the series peak heap, streaming peak heap, materializing (≤ 0.1.3)
    500 K 3.7 MiB 61 MiB
    2 M 3.8 MiB 293 MiB
    8 M 3.9 MiB 1097 MiB

    Why the old path was so expensive: a decoded value is a 64-byte struct against ~2–8 compressed bytes on disk, so materializing a series costs roughly 8–32× its on-disk size (more, transiently, while the slice grows). Versions 0.1.2–0.1.3 held one whole series at a time, which is fine until a single series is hundreds of GB — then no instance size is large enough. 0.1.4 removes that ceiling.

  • Load adds the chunk buffer: each worker accumulates up to --chunk-bytes of raw line protocol before flushing, so migration-host RAM is roughly workers × chunk-bytes (e.g. 4 × 450 MB ≈ 1.8 GB). This dominates the extraction cost, and it is the only migration-host knob that matters. Lower --chunk-bytes and/or --workers to reduce it; both are safe to change (resume/correctness are unaffected).

--start / --end also bound work, not just output: blocks whose time range falls outside the window are skipped straight from the TSM index and never read or decoded. (Before 0.1.4 they filtered output only, after full decode.)

File descriptors. Streaming keeps a handle open per TSM file holding the series being merged. Files whose time ranges cannot interleave are merged as separate passes, so a key split across many files costs one or two handles, not one per file. On a shard whose files all span the same time range, budget workers × files-per-shard descriptors and raise ulimit -n if needed.

On the Arc node, each in-flight import is buffered server-side (~chunk-bytes decompressed + parsed records), so its peak is workers × ~1–1.3 GB — usually the binding constraint (see above).

Always pass --waldir

InfluxDB does not flush the write-ahead log to TSM on shutdown — small or recently-written shards can live entirely in .wal files. On a cold/unmounted volume this is common. When --waldir is given, tsm2arc discovers and migrates WAL-only shards (a shard with no .tsm but non-empty .wal). Without --waldir, it sees only .tsm files, so a WAL-only shard has nothing to find and its data never reaches Arc. (For 2.x, --waldir is auto-detected from engine/wal.)

With --waldir, tsm2arc reads both sources and field-rejoins them per shard. If a point exists in both a TSM file and the WAL (a partially-compacted shard), the WAL value wins (last-write-wins, matching InfluxDB and Arc compaction). The WAL directory mirrors the data directory layout: <waldir>/<db>/<rp>/<shard>/*.wal.

Each source InfluxDB database is routed to the Arc database of the same name (override with --db-map). Data is sent in gzipped chunks bounded at --chunk-bytes of raw line protocol — Arc's import endpoint caps requests at 500 MB of decompressed LP, so the default 450 MB leaves headroom. Transient failures (429, 5xx, network) are retried with exponential backoff; 4xx errors are permanent and abort the run.

Measurement names Arc rejects (dots etc.)

Arc only accepts measurement names matching ^[a-zA-Z][a-zA-Z0-9_-]*$ — the dot is Arc's database.measurement separator in the query layer and in RBAC grant keys, so it can't appear inside a measurement name. InfluxDB is far more permissive (dotted <env>.<service> names are common), so a 1.x/2.x dataset can be full of names Arc will reject with a 400.

tsm2arc validates every name client-side, before anything is sent:

  • --dry-run lists every measurement that would be renamed, skipped, or would abort a load — with point counts — before any network traffic. Start here.
  • --measurement-map old=new (repeatable) and --measurement-map-file PATH (one old=new per line, # comments) rename measurements explicitly. This is the recommended path: you author deterministic names traceable back to source. Map targets are validated at startup; the map may also rename names that are already valid.
  • --on-invalid-measurement controls what happens to a name that is still invalid after the map:
    • fail (default) — abort immediately with an actionable error, before the point is sent. No more dying at chunk 73 on an Arc 400.
    • skip — drop the measurement's points, keep loading, and report exactly what was skipped (names + point counts).
    • map — auto-rename deterministically: every disallowed character becomes _, and a name not starting with a letter gets an m_ prefix (e.g. edge-prod.gateway_services → edge-prod_gateway_services). Distinct source names can collide after sanitizing (a.b and a_b both → a_b), which would merge those measurements — prefer an explicit map when names are close together.

Nothing is silent. Every rename (explicit or auto) and every skip is recorded in the checkpoint database (table measurement_actions: source db, shard, source name, final name, origin, point count) and summarized at the end of the run — so each rename is auditable and reversible, and skipped data is on record rather than quietly missing.

⚠️ Hyphenated names need Arc ≥ 26.09.1 to query. A name like has-hyphen is valid — Arc accepts it at write time and tsm2arc migrates it — but Arc versions before 26.09.1 cannot reference it in SQL at all: the query rewriter didn't resolve quoted identifiers, and an unquoted hyphen parses as subtraction, so no syntax reached the table. From 26.09.1 on, query it quoted — FROM "has-hyphen" (with the x-arc-database header) or FROM "db"."has-hyphen". Unquoted hyphenated names are a SQL parse error on every version — that's SQL grammar, not Arc. If your target Arc predates 26.09.1 and can't be upgraded first, rename at migration time: --measurement-map 'has-hyphen=has_hyphen'. Data migrated with hyphens before an upgrade is stored correctly and becomes queryable as soon as Arc is upgraded — no re-migration needed.

# preview what would happen
tsm2arc --datadir /mnt/influx/data --waldir /mnt/influx/wal --dry-run

# author renames for the dotted names it reported, then load
tsm2arc ... \
  --measurement-map-file renames.map \
  --on-invalid-measurement=fail        # fail if anything is still unmapped

Crash-safe resume

A load is resumable. Progress is tracked per shard — keyed on the stable source ID (1.x database name / 2.x bucket ID) — in a SQLite checkpoint file (--checkpoint, default tsm2arc.checkpoint.db). Each chunk's progress is committed only after Arc returns 2xx (and Arc's import handler flushes to storage before returning, so 2xx means durably persisted).

If a migration is interrupted — process killed, network drop, host reboot — just re-run the exact same command. tsm2arc:

  • skips any shard already fully migrated (no re-extraction),
  • for a partially-migrated shard, seeks straight to where it left off: the checkpoint stores a cursor (series + timestamp of the last acknowledged line), so series before it are never read, already-sent TSM blocks are skipped at the index without being decoded, and sending resumes from the first un-acknowledged chunk within seconds to minutes.

A checkpoint written by tsm2arc ≤ 0.1.4 has no cursor; those shards resume the old way (re-derive deterministically, skip already-sent chunks without re-sending — the heartbeat shows the catch-up as +N skipped on resume). After one 0.1.5 commit the shard has a cursor and later resumes seek.

Because chunk boundaries are a deterministic function of a shard's extraction order and --chunk-bytes, a resumed shard produces byte-identical chunks, so the skip is exact. The only duplication window is a crash between Arc persisting a chunk and tsm2arc recording it: on resume that single chunk is re-sent. Arc compaction collapses the duplicate for tag-bearing series; tagless series retain at most one chunk of duplicate rows per shard per crash (see docs/DESIGN.md §6). A clean (uninterrupted) run produces zero duplicates.

The checkpoint file is safe to keep between runs and is how resume works — keep it alongside the migration. Delete it only to force a full re-migration from scratch.

Resume requires the same shaping flags. The checkpoint records a fingerprint of --chunk-bytes, --start, --end, --db-map, --precision, and — when used — --measurement-map/--measurement-map-file and --on-invalid-measurement (renames and skips change chunk bytes). Resuming with any of these changed would misalign chunk boundaries, so tsm2arc refuses it with a clear error rather than corrupting the migration. To change a shaping flag, start a fresh --checkpoint (a full re-migration). Checkpoints created by tsm2arc ≤ 0.1.2 resume unchanged as long as the new flags stay at their defaults.

InfluxDB 3 (Core/Enterprise) sources — experimental

tsm2arc can read InfluxDB 3 (3.0+) Parquet-engine object stores on local disk (--object-store file, or any store synced to a directory) or in place on S3 with --datadir s3://bucket/prefix — listings and range reads, no scratch-volume sync. Point --datadir at the store root (the directory/prefix holding the node prefix) — the layout is auto-detected, each live table migrates as one unit, and everything else (chunking, resume, --workers, --inflight, --analyze, --redact, measurement policies) works as for 1.x/2.x.

tsm2arc --datadir /mnt/influxdb3-store --arc-url https://arc.example.net --dry-run
# or straight from S3 (credentials from the standard AWS chain;
# --s3-endpoint for MinIO/on-prem gateways):
tsm2arc --datadir s3://my-bucket/influxdb3 --arc-url https://arc.example.net --dry-run

What to know before running:

  • Support tiers (each behavior validated against real stores from live licensed servers):
    • Core 3.0–3.11 (Parquet engine): fully supported.
    • Enterprise (Parquet engine): supported when the compactor has not run — Enterprise cluster layouts (catalog under the cluster prefix) resolve automatically. When compactor state exists (cs//cd//c/), compacted data lives behind a proprietary index and the original gen1 files get deleted, so tsm2arc refuses rather than migrating holes; the supported recipes (nodes in --mode ingest,query, migrate before compaction, or query-export) are printed with the refusal.
    • Pacha-tree engine (.pt, Enterprise 3.11+ new clusters): out of scope; detected and explained — migrate the retained parquet before cleanup-parquet, or export via query from the running server.
  • The WAL is read natively. InfluxDB 3 keeps up to ~10 minutes of the newest writes only in its WAL, and a clean shutdown does not flush them (no snapshot on shutdown). tsm2arc decodes un-snapshotted WAL files in-process (the upstream codec compiled to WebAssembly, embedded in the binary — still a single static Go binary) and merges those rows into the migration, so nothing is left behind; --skip-wal skips the decode explicitly.
  • Names come from the catalog — every era. JSON catalogs (3.0–3.9) and the binary 3.10+/3.11 catalog both resolve automatically, including the mandatory log replay past the last checkpoint (catalogs checkpoint lazily; recent databases/tables often exist only in logs). --v3-db / --v3-table remain as manual overrides; unresolved ids fail loudly, never guess.
  • Duplicates resolve last-write-wins deterministically (the same overwrite semantics InfluxDB 3 applies at query time), and emission order is a pure function of the store, so resume is byte-exact — the live file set is part of the checkpoint fingerprint, and a store that changed under a checkpoint fails as "different settings".
  • Not yet: multi-node (Enterprise cluster) stores, GCS/Azure sources, Enterprise-with-compaction validation. Tracking: #10.

Validate against a local InfluxDB

cd fixture
docker compose up -d
./seed.sh                       # writes a known dataset (all field types,
                                # tagless measurement, pre-epoch timestamp,
                                # two databases)
docker compose restart influxdb # clean shutdown flushes WAL → TSM

# extract and compare against the oracle printed by seed.sh
../tsm2arc --datadir ./data/influxdb/data --waldir ./data/influxdb/wal \
           --dry-run --sample 20

Design

Full design + the verified Arc ingest constraints (500 MB import cap on decompressed bytes, admin auth, append-only ingest with compaction-time dedupe, resume protocol) are in docs/DESIGN.md.

How extraction works

InfluxDB stores each field of a point as a separate TSM key (cpu,host=a#!~#usage and cpu,host=a#!~#cores), each with its own timestamp+value stream. tsm2arc:

  1. parses the TSM index (header/index/footer) for each file;
  2. decodes every block via native Go codec implementations;
  3. rejoins fields by (series, timestamp) so multi-field points are reconstructed as single line-protocol lines;
  4. emits line protocol with original nanosecond timestamps (pre-1970 supported), routed to the Arc database matching the source InfluxDB database.

Development

go test ./...          # includes round-trip tests vs the real influx encoder
go vet ./...
gofmt -l .

github.com/influxdata/influxdb v1.7.11 is a test-only dependency (the oracle for codec validation); the production binary does not link it.

About

Migrate InfluxDB 1.x & 2.x data into Arc by reading TSM/WAL files directly off disk — no running influxd required.

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages