Migrate InfluxDB 1.x (1.7/1.8) and 2.x (2.0–2.7) data into
Arc by reading TSM and WAL files
directly off disk — no running influxd required. InfluxDB 3
(Parquet engine) sources are supported experimentally — see
InfluxDB 3 sources.
The on-disk TSM/WAL format is the same across 1.x and 2.x; tsm2arc auto-detects
the layout. For 2.x it resolves bucket IDs to readable names from influxd.bolt
and skips InfluxDB's internal system buckets (e.g. _monitoring, _tasks).
Built for the case where InfluxDB data sits on cold/unmounted volumes (e.g. EBS
snapshots) that can be mounted read-only but are not served by any InfluxDB
instance. tsm2arc decodes the TSM/WAL block codecs natively in Go, reconstructs
multi-field line-protocol points, and streams them into Arc's
/api/v1/import/lp endpoint in resumable, size-bounded, parallel chunks.
Reusable beyond Arc. The extraction side produces standard InfluxDB line protocol; the Arc sink is just the first sink. The TSM/WAL decoder is Apache-2.0 and can be reused to migrate InfluxDB 1.x data into other systems — contributions of new sinks (ClickHouse, QuestDB, TimescaleDB, …) are welcome. See CONTRIBUTING.md.
Download a prebuilt binary from Releases (Linux/macOS/Windows, amd64/arm64), or:
# from source (Go 1.25+)
go install github.com/basekick-labs/tsm2arc/cmd/tsm2arc@latest
# container
docker run --rm ghcr.io/basekick-labs/tsm2arc:latest --versionEach release ships an SBOM (SPDX) and checksums.txt; the container image is
multi-arch (linux amd64/arm64) on GHCR.
Feature-complete. Capabilities:
| Capability | State |
|---|---|
Native TSM reader + field-rejoin + LP encode + --dry-run |
✅ |
Chunked gzip POST to Arc /api/v1/import/lp (per-DB routing) |
✅ |
| SQLite checkpoint + crash-safe resume | ✅ |
WAL (.wal) reader — merged with TSM per shard |
✅ |
Parallel workers (--workers) + live progress reporting |
✅ |
| InfluxDB 2.x layout auto-detection + bucket-name resolution | ✅ |
Measurement rename map + invalid-name policy (fail/skip/map) with a checkpoint audit trail |
✅ |
| InfluxDB 3.x Parquet-engine object stores (local disk) — see InfluxDB 3 sources | 🧪 experimental |
The TSM/WAL codecs (timestamp, float, integer, unsigned, boolean, string) are
validated against the real InfluxDB 1.7.11 encoder in unit tests
(internal/tsm/decode_test.go, file_test.go, internal/wal/wal_test.go) and
cross-checked against real InfluxDB 2.7 data.
go build ./cmd/tsm2arc
# InfluxDB 1.x — point at the data dir (or its parent; layout is auto-detected)
./tsm2arc --datadir /var/lib/influxdb --dry-run --sample 10
# InfluxDB 2.x — point at the v2 root (~/.influxdbv2 or the mounted volume);
# engine/data, engine/wal, and influxd.bolt (for bucket names) are auto-detected
./tsm2arc --datadir /var/lib/influxdb2 --dry-run --sample 10Dry-run discovers shards, decodes every block, reconstructs points, and prints per-database/bucket counts + sample line protocol — without writing to Arc. This is the safe first contact with the source data.
tsm2arc auto-detects 1.x vs 2.x. You can point --datadir at the version's
natural root and it resolves the rest:
- 1.x: the InfluxDB root (containing
data/), or…/datadirectly. - 2.x: the v2 root (containing
engine/andinfluxd.bolt),…/engine, or…/engine/datadirectly. Bucket names come frominfluxd.bolt; override its location with--boltif it lives elsewhere. Without it, buckets fall back to their 16-hex IDs as database names.
./tsm2arc \
--datadir /mnt/influxdb/data \
--waldir /mnt/influxdb/wal \ # IMPORTANT: include the WAL (see below)
--arc-url https://arc.example.net \
--token "$ARC_TOKEN" \ # admin-tier token (or ARC_TOKEN env)
--verbose
# options:
# --waldir DIR InfluxDB WAL directory — read un-flushed data too
# (auto-detected for 2.x as engine/wal)
# --bolt PATH InfluxDB 2.x influxd.bolt for bucket names (auto-detected)
# --workers N concurrent shards to migrate (default 2)
# --db-map old=new rename a source DB/bucket to a different Arc DB (repeatable)
# --database-filter db migrate only this source DB/bucket (repeatable)
# --measurement-map old=new rename a source measurement (repeatable; new must
# satisfy Arc's name rule — see below)
# --measurement-map-file PATH file of measurement renames, one old=new per line
# (blank lines and # comments ignored)
# --on-invalid-measurement MODE what to do with a measurement name Arc would
# reject, after --measurement-map is applied:
# fail (default) | skip (drop + report) | map
# (deterministic auto-rename)
# --chunk-bytes SIZE raw-LP per import request; bytes or a suffix like
# 450MB (must be <500MB; default 450MB)
# --checkpoint PATH SQLite resume store (default tsm2arc.checkpoint.db)
# --start / --end RFC3339 UTC time filters
# --precision ns|us|ms|s precision value sent to Arc (default ns; tsm2arc always emits ns)
# --include-internal also migrate InfluxDB 1.x's _internal database
# (2.x system buckets _monitoring/_tasks are always skipped)
# --inflight N concurrent import POSTs per shard (default 1);
# commits stay strictly ordered, so resume is exact
# at any value. >1 widens the tagless crash-duplicate
# bound to <=N chunks/shard and multiplies host and
# Arc memory (see the runbook)
# --pipeline overlap extraction with upload (default true; =false
# reverts to serial send and saves one chunk buffer
# of memory per worker)
# --shard-split N max concurrent merge tasks per shard (default 1).
# Output is byte-identical to serial, so resume works
# across different values. Requires --merge-memory
# --merge-memory SIZE per-shard admission budget for concurrent merges
# (e.g. 24GB). Concurrent merges are bounded by
# memory, not cores: a merge holds ~one decoded block
# per (file x field) stream. Oversized tasks run alone
# --index-cache SIZE per-shard budget for cached TSM file indexes
# (default 2GB; 0 disables). Avoids re-parsing every
# file's index once per series during extraction —
# a large CPU cost on shards with many series.
# --analyze index-only shard profile (fast; nothing decoded or
# sent): series/file/key counts per shard and window-
# split profiles for the largest merge runs
# --analyze-runs N text output: show the N largest runs per shard
# (default 5, labeled when it truncates; 0 = all)
# --format text|json with --analyze: json emits EVERY run of every shard
# as one JSON document on a pure stdout (for tooling)
# --send-timeout DUR per-attempt deadline for one import POST (default
# auto: 2m + 1s per MiB of --chunk-bytes). Bounds dead
# connections; too low re-sends healthy imports
# --stall-warn DUR warn when a shard makes no progress for this long
# (default auto: send-timeout + 2m, min 5m)
# --redact with --analyze: replace database, retention policy,
# and series names with stable hashed identifiers so
# the report can be shared outside your organization
# --dry-run extract + count, do not write to Arc
# --sample N print N sample LP lines per DB in --dry-run (default 5)
# --verbose per-shard / per-chunk logging
# --version print version and exit--workers N migrates N shards concurrently. Shards are fully independent (each
has its own chunk sequence and checkpoint rows), so this scales cleanly. A live
heartbeat reports shards-done, chunks, rows, MB, and throughput as the migration
runs.
Size --workers against the Arc node's memory, not the migration host's:
Arc's import endpoint buffers each request fully in memory (~chunk-bytes
decompressed + parsed records), so peak transient Arc-side memory is roughly
workers × ~1–1.3 GB at the default 450 MB chunk size. The default of 2 is
conservative; raise it deliberately if the Arc node has headroom (the customer's
big dedicated migration host is rarely the bottleneck — Arc is). Resume and
correctness are unaffected by the worker count.
tsm2arc extracts by streaming one series at a time and, within a series, one TSM block at a time. It indexes only the key list up front, then merges each series straight off lazy block cursors. So on the migration host:
-
Extraction /
--dry-runis near-constant memory — it does not scale with the shard, the dataset, or the largest series. Measured on one series written as 1000-value blocks:values in the series peak heap, streaming peak heap, materializing (≤ 0.1.3) 500 K 3.7 MiB 61 MiB 2 M 3.8 MiB 293 MiB 8 M 3.9 MiB 1097 MiB Why the old path was so expensive: a decoded value is a 64-byte struct against ~2–8 compressed bytes on disk, so materializing a series costs roughly 8–32× its on-disk size (more, transiently, while the slice grows). Versions 0.1.2–0.1.3 held one whole series at a time, which is fine until a single series is hundreds of GB — then no instance size is large enough. 0.1.4 removes that ceiling.
-
Load adds the chunk buffer: each worker accumulates up to
--chunk-bytesof raw line protocol before flushing, so migration-host RAM is roughlyworkers × chunk-bytes(e.g.4 × 450 MB ≈ 1.8 GB). This dominates the extraction cost, and it is the only migration-host knob that matters. Lower--chunk-bytesand/or--workersto reduce it; both are safe to change (resume/correctness are unaffected).
--start / --end also bound work, not just output: blocks whose time range
falls outside the window are skipped straight from the TSM index and never read
or decoded. (Before 0.1.4 they filtered output only, after full decode.)
File descriptors. Streaming keeps a handle open per TSM file holding the
series being merged. Files whose time ranges cannot interleave are merged as
separate passes, so a key split across many files costs one or two handles, not
one per file. On a shard whose files all span the same time range, budget
workers × files-per-shard descriptors and raise ulimit -n if needed.
On the Arc node, each in-flight import is buffered server-side (~chunk-bytes
decompressed + parsed records), so its peak is workers × ~1–1.3 GB — usually
the binding constraint (see above).
InfluxDB does not flush the write-ahead log to TSM on shutdown — small or
recently-written shards can live entirely in .wal files. On a cold/unmounted
volume this is common. When --waldir is given, tsm2arc discovers and migrates
WAL-only shards (a shard with no .tsm but non-empty .wal). Without
--waldir, it sees only .tsm files, so a WAL-only shard has nothing to find
and its data never reaches Arc. (For 2.x, --waldir is auto-detected from
engine/wal.)
With --waldir, tsm2arc reads both sources and field-rejoins them per shard. If a
point exists in both a TSM file and the WAL (a partially-compacted shard), the WAL
value wins (last-write-wins, matching InfluxDB and Arc compaction). The WAL
directory mirrors the data directory layout: <waldir>/<db>/<rp>/<shard>/*.wal.
Each source InfluxDB database is routed to the Arc database of the same name
(override with --db-map). Data is sent in gzipped chunks bounded at
--chunk-bytes of raw line protocol — Arc's import endpoint caps requests at
500 MB of decompressed LP, so the default 450 MB leaves headroom. Transient
failures (429, 5xx, network) are retried with exponential backoff; 4xx errors
are permanent and abort the run.
Arc only accepts measurement names matching ^[a-zA-Z][a-zA-Z0-9_-]*$ — the
dot is Arc's database.measurement separator in the query layer and in RBAC
grant keys, so it can't appear inside a measurement name. InfluxDB is far more
permissive (dotted <env>.<service> names are common), so a 1.x/2.x dataset
can be full of names Arc will reject with a 400.
tsm2arc validates every name client-side, before anything is sent:
--dry-runlists every measurement that would be renamed, skipped, or would abort a load — with point counts — before any network traffic. Start here.--measurement-map old=new(repeatable) and--measurement-map-file PATH(oneold=newper line,#comments) rename measurements explicitly. This is the recommended path: you author deterministic names traceable back to source. Map targets are validated at startup; the map may also rename names that are already valid.--on-invalid-measurementcontrols what happens to a name that is still invalid after the map:fail(default) — abort immediately with an actionable error, before the point is sent. No more dying at chunk 73 on an Arc 400.skip— drop the measurement's points, keep loading, and report exactly what was skipped (names + point counts).map— auto-rename deterministically: every disallowed character becomes_, and a name not starting with a letter gets anm_prefix (e.g.edge-prod.gateway_services→edge-prod_gateway_services). Distinct source names can collide after sanitizing (a.banda_bboth →a_b), which would merge those measurements — prefer an explicit map when names are close together.
Nothing is silent. Every rename (explicit or auto) and every skip is
recorded in the checkpoint database (table measurement_actions: source db,
shard, source name, final name, origin, point count) and summarized at the end
of the run — so each rename is auditable and reversible, and skipped data is
on record rather than quietly missing.
⚠️ Hyphenated names need Arc ≥ 26.09.1 to query. A name likehas-hyphenis valid — Arc accepts it at write time and tsm2arc migrates it — but Arc versions before 26.09.1 cannot reference it in SQL at all: the query rewriter didn't resolve quoted identifiers, and an unquoted hyphen parses as subtraction, so no syntax reached the table. From 26.09.1 on, query it quoted —FROM "has-hyphen"(with thex-arc-databaseheader) orFROM "db"."has-hyphen". Unquoted hyphenated names are a SQL parse error on every version — that's SQL grammar, not Arc. If your target Arc predates 26.09.1 and can't be upgraded first, rename at migration time:--measurement-map 'has-hyphen=has_hyphen'. Data migrated with hyphens before an upgrade is stored correctly and becomes queryable as soon as Arc is upgraded — no re-migration needed.
# preview what would happen
tsm2arc --datadir /mnt/influx/data --waldir /mnt/influx/wal --dry-run
# author renames for the dotted names it reported, then load
tsm2arc ... \
--measurement-map-file renames.map \
--on-invalid-measurement=fail # fail if anything is still unmappedA load is resumable. Progress is tracked per shard — keyed on the stable
source ID (1.x database name / 2.x bucket ID) — in a SQLite checkpoint file
(--checkpoint, default tsm2arc.checkpoint.db). Each chunk's progress is
committed only after Arc returns 2xx (and Arc's import handler flushes to
storage before returning, so 2xx means durably persisted).
If a migration is interrupted — process killed, network drop, host reboot — just re-run the exact same command. tsm2arc:
- skips any shard already fully migrated (no re-extraction),
- for a partially-migrated shard, seeks straight to where it left off: the checkpoint stores a cursor (series + timestamp of the last acknowledged line), so series before it are never read, already-sent TSM blocks are skipped at the index without being decoded, and sending resumes from the first un-acknowledged chunk within seconds to minutes.
A checkpoint written by tsm2arc ≤ 0.1.4 has no cursor; those shards resume the
old way (re-derive deterministically, skip already-sent chunks without
re-sending — the heartbeat shows the catch-up as +N skipped on resume). After
one 0.1.5 commit the shard has a cursor and later resumes seek.
Because chunk boundaries are a deterministic function of a shard's extraction
order and --chunk-bytes, a resumed shard produces byte-identical chunks, so the
skip is exact. The only duplication window is a crash between Arc persisting a
chunk and tsm2arc recording it: on resume that single chunk is re-sent. Arc
compaction collapses the duplicate for tag-bearing series; tagless series retain
at most one chunk of duplicate rows per shard per crash (see
docs/DESIGN.md §6). A clean (uninterrupted) run produces zero
duplicates.
The checkpoint file is safe to keep between runs and is how resume works — keep it alongside the migration. Delete it only to force a full re-migration from scratch.
Resume requires the same shaping flags. The checkpoint records a fingerprint
of --chunk-bytes, --start, --end, --db-map, --precision, and — when
used — --measurement-map/--measurement-map-file and
--on-invalid-measurement (renames and skips change chunk bytes). Resuming
with any of these changed would misalign chunk boundaries, so tsm2arc refuses
it with a clear error rather than corrupting the migration. To change a shaping
flag, start a fresh --checkpoint (a full re-migration). Checkpoints created
by tsm2arc ≤ 0.1.2 resume unchanged as long as the new flags stay at their
defaults.
tsm2arc can read InfluxDB 3 (3.0+) Parquet-engine object stores on local
disk (--object-store file, or any store synced to a directory) or in
place on S3 with --datadir s3://bucket/prefix — listings and range reads,
no scratch-volume sync. Point --datadir at the store root (the
directory/prefix holding the node prefix) — the layout is auto-detected, each
live table migrates as one unit, and everything else (chunking, resume,
--workers, --inflight, --analyze, --redact, measurement policies)
works as for 1.x/2.x.
tsm2arc --datadir /mnt/influxdb3-store --arc-url https://arc.example.net --dry-run
# or straight from S3 (credentials from the standard AWS chain;
# --s3-endpoint for MinIO/on-prem gateways):
tsm2arc --datadir s3://my-bucket/influxdb3 --arc-url https://arc.example.net --dry-runWhat to know before running:
- Support tiers (each behavior validated against real stores from live
licensed servers):
- Core 3.0–3.11 (Parquet engine): fully supported.
- Enterprise (Parquet engine): supported when the compactor has not
run — Enterprise cluster layouts (catalog under the cluster prefix)
resolve automatically. When compactor state exists (
cs//cd//c/), compacted data lives behind a proprietary index and the original gen1 files get deleted, so tsm2arc refuses rather than migrating holes; the supported recipes (nodes in--mode ingest,query, migrate before compaction, or query-export) are printed with the refusal. - Pacha-tree engine (
.pt, Enterprise 3.11+ new clusters): out of scope; detected and explained — migrate the retained parquet beforecleanup-parquet, or export via query from the running server.
- The WAL is read natively. InfluxDB 3 keeps up to ~10 minutes of the
newest writes only in its WAL, and a clean shutdown does not flush them
(no snapshot on shutdown). tsm2arc decodes un-snapshotted WAL files
in-process (the upstream codec compiled to WebAssembly, embedded in the
binary — still a single static Go binary) and merges those rows into the
migration, so nothing is left behind;
--skip-walskips the decode explicitly. - Names come from the catalog — every era. JSON catalogs (3.0–3.9) and
the binary 3.10+/3.11 catalog both resolve automatically, including the
mandatory log replay past the last checkpoint (catalogs checkpoint lazily;
recent databases/tables often exist only in logs).
--v3-db/--v3-tableremain as manual overrides; unresolved ids fail loudly, never guess. - Duplicates resolve last-write-wins deterministically (the same overwrite semantics InfluxDB 3 applies at query time), and emission order is a pure function of the store, so resume is byte-exact — the live file set is part of the checkpoint fingerprint, and a store that changed under a checkpoint fails as "different settings".
- Not yet: multi-node (Enterprise cluster) stores, GCS/Azure sources, Enterprise-with-compaction validation. Tracking: #10.
cd fixture
docker compose up -d
./seed.sh # writes a known dataset (all field types,
# tagless measurement, pre-epoch timestamp,
# two databases)
docker compose restart influxdb # clean shutdown flushes WAL → TSM
# extract and compare against the oracle printed by seed.sh
../tsm2arc --datadir ./data/influxdb/data --waldir ./data/influxdb/wal \
--dry-run --sample 20Full design + the verified Arc ingest constraints (500 MB import cap on decompressed bytes, admin auth, append-only ingest with compaction-time dedupe, resume protocol) are in docs/DESIGN.md.
InfluxDB stores each field of a point as a separate TSM key
(cpu,host=a#!~#usage and cpu,host=a#!~#cores), each with its own
timestamp+value stream. tsm2arc:
- parses the TSM index (header/index/footer) for each file;
- decodes every block via native Go codec implementations;
- rejoins fields by (series, timestamp) so multi-field points are reconstructed as single line-protocol lines;
- emits line protocol with original nanosecond timestamps (pre-1970 supported), routed to the Arc database matching the source InfluxDB database.
go test ./... # includes round-trip tests vs the real influx encoder
go vet ./...
gofmt -l .github.com/influxdata/influxdb v1.7.11 is a test-only dependency (the
oracle for codec validation); the production binary does not link it.