Skip to content

Allow constructing a Matcher from a caller-supplied corpus #59

Description

@andrew

brief currently uses the github.com/git-pkgs/licensecheck fork for one job: read a LICENSE file and return an SPDX identifier (detect.detectLicenseType). Moving that call to github.com/git-pkgs/licenses would let us retire the fork, but importing this package pulls in the full embedded ScanCode corpus.

Measured on darwin/arm64, go1.26.7:

binary size
brief with git-pkgs/licensecheck v0.4.1 48.8 MB
brief with git-pkgs/licenses v0.8.0 added 60.5 MB

The delta is the 11.5 MB internal/corpus/corpus.bin.gz embed. brief already depends on git-pkgs/magic, git-pkgs/manifests and git-pkgs/spdx, so the transitive graph adds nothing else.

The only public constructor is licenses.New, which unconditionally calls internal/corpus.Load, which references the //go:embed var. There is no way to construct a Matcher against a smaller corpus, and internal/corpus (which has Read/Write for the on-disk format) is not importable from outside the module.

For a LICENSE-file classifier, only a fraction of the corpus is relevant. Rule counts in the v0.8.0 embed (ScanCode 33.0.0rc1, 39,215 rules total):

flag rules
FlagLicenseText 6,959
FlagLicenseNotice 14,917
FlagLicenseReference 11,893
FlagLicenseTag 3,853
other (intro/clue/false-positive) 1,593

Proposal:

  1. Add func NewFromReader(r io.Reader, options ...Option) (*Matcher, error) that decodes a corpus with internal/corpus.Read and builds a matchEngine from it, bypassing the embedded index. The Go linker drops an unreferenced //go:embed package var (verified: a 5 MB embed referenced only by an unreachable function is not linked), so a caller that uses only NewFromReader does not carry the 11.5 MB.
  2. Teach the corpus builder to emit a filtered index, e.g. --rule-flags text or --rule-flags text,notice, which writes only the selected rules and rebuilds the Aho-Corasick automaton and vocabulary over that subset. internal/corpus.Write already serialises an arbitrary Index; the missing piece is regenerating Index.Automaton and trimming Index.Vocabulary/StopwordIDs for the reduced rule set. Callers embed the resulting corpus.bin.gz themselves and pass it to NewFromReader.

licensecheck's builtin.dfa is 2.9 MB, so a text-rules-only corpus in the same ballpark would let brief swap without growing, and dependents/downstream (which pull licensecheck in transitively via brief) would follow.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions