Skip to content

archive_otu_table: Add version 5, a more compact on-disk encoding - #316

Merged
wwood merged 8 commits into
mainfrom
claude/frame-shift-chunk-processing-icz3u2
Aug 11, 2026
Merged

wwood merged 8 commits into
mainfrom
claude/frame-shift-chunk-processing-icz3u2

Conversation

@wwood

@wwood wwood commented Aug 6, 2026

Copy link
Copy Markdown
Owner

Archive OTU tables repeat themselves heavily: every OTU carries the sample
name and taxonomy assignment method again, the field names are implicit in
row order so nothing groups like values together, and a read hitting two
markers has its sequence stored once per marker. Version 5 removes those
three kinds of repetition without changing what the table holds.

  • otus is stored column-wise rather than as a list of rows.
  • Fields whose value is identical in every OTU move to constant_fields,
    with num_otus recording the row count that the columns then no longer
    give when every field happens to be constant.
  • Read sequences are dereplicated into a shared reads list.

Reads are referred to by 0 for "the next read not yet used" and by n for a
repeat of reads[n-1], rather than by plain position. Plain positions cost
more than they save - they are mostly-ascending integers, which compresses
badly, in place of duplicate reads that the compressor was already matching
nearly for free; measured, they took gzip -9 from 36.6 KB to 38.2 KB. The
scheme used here leaves the common case a run of zeros instead.

Versions 1-4 are still read, and ArchiveOtuTable presents the same list of
rows whichever version it read, so nothing downstream of read() changes.
Code that parses the JSON itself does need updating, which is why the GroopM
API stability test now decodes the version 5 layout directly rather than
going through ArchiveOtuTable and hiding the difference.

The gain is modest: gzip -9 36.6 -> 35.6 KB and zstd -19 32.1 -> 30.4 KB on
test/data/SRR8653040.json. What is left is mostly the read sequences plus
the per-package sha256s, which are incompressible and are 13% of that
gzipped file on their own. Storing fewer or shorter reads is the only thing
that would shrink an archive substantially.

Co-Authored-By: Claude Opus 5 noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01SxX9hkwfCutcgmNgHHYffo

Archive OTU tables repeat themselves heavily: every OTU carries the sample
name and taxonomy assignment method again, the field names are implicit in
row order so nothing groups like values together, and a read hitting two
markers has its sequence stored once per marker. Version 5 removes those
three kinds of repetition without changing what the table holds.

  * otus is stored column-wise rather than as a list of rows.
  * Fields whose value is identical in every OTU move to constant_fields,
    with num_otus recording the row count that the columns then no longer
    give when every field happens to be constant.
  * Read sequences are dereplicated into a shared reads list.

Reads are referred to by 0 for "the next read not yet used" and by n for a
repeat of reads[n-1], rather than by plain position. Plain positions cost
more than they save - they are mostly-ascending integers, which compresses
badly, in place of duplicate reads that the compressor was already matching
nearly for free; measured, they took gzip -9 from 36.6 KB to 38.2 KB. The
scheme used here leaves the common case a run of zeros instead.

Versions 1-4 are still read, and ArchiveOtuTable presents the same list of
rows whichever version it read, so nothing downstream of read() changes.
Code that parses the JSON itself does need updating, which is why the GroopM
API stability test now decodes the version 5 layout directly rather than
going through ArchiveOtuTable and hiding the difference.

The gain is modest: gzip -9 36.6 -> 35.6 KB and zstd -19 32.1 -> 30.4 KB on
test/data/SRR8653040.json. What is left is mostly the read sequences plus
the per-package sha256s, which are incompressible and are 13% of that
gzipped file on their own. Storing fewer or shorter reads is the only thing
that would shrink an archive substantially.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxX9hkwfCutcgmNgHHYffo

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: be0faf11f4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


class ArchiveOtuTable:
version = 4
version = 5

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Allow merging version 4 and 5 archives

When summarise --collapse-paired-with-unpaired or --collapse-to-sample-name receives a pre-upgrade v4 archive together with a newly generated v5 archive, Summariser.write_collapsed_paired_with_unpaired_otu_table() rejects them at singlem/summariser.py:384-385 with "Version mismatch between archives". Since v4 and v5 have identical fields and ArchiveOtuTable.read() normalizes both to rows, changing the writer default here makes previously valid merge workflows fail unless the merge path permits these two versions.

Useful? React with 👍 / 👎.


class ArchiveOtuTable:
version = 4
version = 5

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the enabled Annoy archive assertion

When annoy is installed, test_annoy_only_assignment_output_archive_single is enabled and still compares the command output against an exact version-4 row-oriented JSON object at test/test_pipe.py:1348-1349; every successful archive produced after this version change is v5 and column-oriented, so that test necessarily fails. Update it to decode through ArchiveOtuTable or assert the v5 layout, as the other archive-output tests now do.

AGENTS.md reference: AGENTS.md:L1-L1

Useful? React with 👍 / 👎.

claude and others added 7 commits August 6, 2026 12:30
Resolving the ambiguous base that repairing a deletion leaves behind is
the one step of pipe whose answer depends on the rest of the dataset: the
base is taken from the most abundant near-identical window, and which
window that is depends on what else was seen. A run over one chunk of a
dataset therefore resolved against that chunk alone, and the chunks no
longer combined into what the whole dataset would have given - a read
could be left ambiguous for want of a donor that was in another chunk, or
resolved to a different donor entirely, splitting counts across sequences
and reaching the taxonomic profile.

So a chunked run now leaves those bases alone, and summarise fills them in
once the chunks are combined, where every chunk is present:

  singlem summarise --collapse-to-sample-name S \
    --input-archive-otu-tables chunk*.json \
    --output-archive-otu-table combined.json \
    --resolve-ambiguous-windows

Both paths go through the new shared resolve_windows(), so they make the
identical choice of donor; pipe weighs abundance by per-read objects and
summarise by num_hits, which are the same numbers. Resolution happens
before the OTUs are grouped, so a window that resolves to one already
present is grouped with it and takes its taxonomy - the taxonomy the whole
dataset would have given it, pipe resolving before assigning taxonomy.

Carrying the provenance that far needs it on disk, so archive OTU table
version 5 gains a reads_with_repaired_deletions field: the indices of the
reads whose window has a base that frameshift repair inserted, as opposed
to one that was already in the raw read. It is None when an OTU has none,
so constant hoisting collapses the column to a single entry on any run
without frameshift repair.

An OTU with only some such reads is left ambiguous and counted in a
warning, since resolving it would mean splitting the OTU and the per-read
coverage needed to divide up its coverage is not in the archive. That
takes two reads with the very same window, one whose ambiguous base came
from a repair and one whose was in the raw read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxX9hkwfCutcgmNgHHYffo
… table

The extended OTU table prints every field the archive has, so adding
reads_with_repaired_deletions gave it a thirteenth column. The field
indexes into an OTU's read names, which is of no use to someone reading a
TSV, so changing that format for it benefits nobody.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxX9hkwfCutcgmNgHHYffo
read_names, read_unaligned_sequences and nucleotides_aligned are parallel
lists in an archive OTU table - index i describes the same read in each,
and renew reads them that way - but add_info sorted only the first two by
read name, leaving the third in collection order.

The two orders only differ when the reads of one OTU do not all have the
same aligned length. That takes an insert column inside the window's span:
_best_position_to_chosen_positions() skips those columns when choosing the
window, while _nucleotide_alignment() counts them towards aligned_length,
so a read with a residue there and a read with a gap there can share a
window and not an aligned length. Rare enough that nucleotides_aligned has
a single distinct value in all 170 OTUs of the two real archives in
test/data, which is why this has not been visible.

Sort the four per-read lists together in one place instead, extracted as
_read_columns_sorted_by_name() so the invariant can be tested directly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxX9hkwfCutcgmNgHHYffo
An archive written before this branch cannot be combined with one written
after it, since read_archive_table() required every input to be the same
version and version 5 is now the default. The versions differ only by fields
appended to the end, and ArchiveOtuTable gives rows either way, so combine on
the wider field list instead: OTUs from the older archive have no value for
the fields its version lacked, and the output is written at whichever version
owns the merged fields. Combining archives that are all of one older version
still gives that version out.

Also fixes --assignment-method annoy failing on single-ended input with
UnboundLocalError on read_name_to_fullseq, which the annoy branch set only
when analysing pairs while the DIAMOND branch below it sets both. Pre-existing
rather than from this branch - confirmed by running the annoy tests at the
commit this work started from - and missed because those tests are skipped
unless annoy is installed, which it is not in CI. Fixing it is what makes the
archive output test below verifiable.

test_annoy_only_assignment_output_archive_single compared the output against
an exact version 4 JSON object, so decode it through ArchiveOtuTable as the
other archive output tests now do.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxX9hkwfCutcgmNgHHYffo
@wwood
wwood merged commit 7662ab6 into main Aug 11, 2026
2 checks passed
@wwood
wwood deleted the claude/frame-shift-chunk-processing-icz3u2 branch August 13, 2026 04:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants