Skip to content

Fix UTF-16 surrogate validation - #291

Open
efegokdemir wants to merge 1 commit into
apache:mainfrom
efegokdemir:codex/issue-287-utf8-surrogate-validation
Open

efegokdemir wants to merge 1 commit into
apache:mainfrom
efegokdemir:codex/issue-287-utf8-surrogate-validation

Conversation

@efegokdemir

@efegokdemir efegokdemir commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #287 by correcting UTF-16 surrogate validation in UTF8Util.isWellFormed.

Problem

The previous logic rejected valid high/low surrogate pairs and accepted unpaired surrogates, including a trailing high surrogate.

Fix

  • Accept valid surrogate pairs.
  • Reject an unpaired low surrogate at the start or in the sequence.
  • Reject a high surrogate not followed by a low surrogate.
  • Reject a trailing high surrogate after the loop.
  • Preserve the existing null-character and sanity checks.

Regression tests

Updated UTF8UtilTest.mustNotChars to cover:

  • a valid surrogate pair;
  • an unpaired high surrogate;
  • an unpaired low surrogate;
  • a trailing high surrogate.

Testing

  • ./mvnw -pl bifromq-util -am -DskipTests=false test — attempted but blocked: the environment has no Java runtime.
  • git diff --check — passed.

Closes #287.

Contribution

AI assistance was used during investigation and implementation; the change and validation results were reviewed before submission.

Signed-off-by: Efe Gökdemir <efe@rexcode.co.uk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] UTF8Util.isWellFormed inverts the surrogate-pair check: valid non-BMP characters are rejected, lone surrogates accepted

1 participant