Skip to content

Bump to Rust 2024 Edition - #102

Closed
wks wants to merge 3793 commits into
mmtk:masterfrom
wks:update/mmtk/edition-2024
Closed

wks wants to merge 3793 commits into
mmtk:masterfrom
wks:update/mmtk/edition-2024

Conversation

@wks

@wks wks commented Sep 28, 2026

Copy link
Copy Markdown

Upstream PR: mmtk/mmtk-core#1596

fingolfin and others added 30 commits September 3, 2026 12:43
`jl_init_options` is normally called by the libjulia loader
(`cli/loader_lib.c`) before any other runtime code. In a static build of
libjulia-internal (JuliaLang#62868) there is no loader, and
`jl_autoinit_and_adopt_thread` — the trampoline that images call on
first entry — becomes the first runtime entry point, so `jl_options`
would never be initialized. Call `jl_init_options` there before
`jl_init_with_image_handle`.

`jl_init_options` guards itself with `jl_options_initialized`, so this
is a no-op for the regular shared build.

Part of the static libjulia-internal work.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Emit load/store pairs for small padding-free pointer-free values,
keeping source and destination alias metadata separate.

Helps quite a bit with vectorization over a variety of common types

```julia
Base.@noinline function _copyto_bench!(dst, src)
    for i in eachindex(dst, src)
        dst[i] = src[i]
    end
    dst
end

n = 1024

T = Tuple{Bool, UInt8}
src = (rand(T, n));
dst = copy(src);

@Btime _copyto_bench!($dst, $src);
```
<img width="866" height="323" alt="image"
src="https://github.com/user-attachments/assets/ff96c25a-d443-4163-8874-bb23688c9ca3"
/>

Fixes JuliaLang#60409

Coauthored-by: Codex (GPT-5.6 Sol xhigh)
…candidates (JuliaLang#62396)

Inspired by JuliaLang#62262, Claude Fable identified a bug in the current
algorithm, as well as proposed some improvements for the accuracy of the
ambiguity computation. This still needs some human intervention to make
sure it does things correctly / efficiently (note that one of the
commits is just NFC code moving).

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`jl_simulate_longjmp` resolves an in-signal throw by decoding the target
`jl_jmp_buf` into the interrupted `ucontext` and letting `sigreturn`
perform the jump. The FreeBSD/AArch64 implementation got this wrong in
two ways, so every throw out of a signal handler—`DivideError`,
`StackOverflowError`, `ReadOnlyMemoryError`, and the `safe_restore`
recovery used by `jl_` and the profiler—crashed the process instead.

First, the jump buffer offsets were off by two. FreeBSD declares the
buffer as `__int128_t _sjb[]` in `<machine/setjmp.h>`, but that element
type only forces the size and alignment; `_setjmp` stores packed 8-byte
words:

    [0] magic, [1] sp, [2..13] x19..x30, [14..21] d8..d15

Decoding from index 0 onward put the `setjmp` magic into `x19`, the
saved `x28` into `gp_lr`, and the saved frame pointer into `gp_sp`. The
`assert(gp_sp % 16 == 0)` sanity check passed regardless, since a frame
pointer is 16-byte aligned too.

Second, the resume pc was never set. Every other architecture ends by
assigning the restored return address to the pc (`mc->pc = mc->regs[30]`
on Linux/AArch64, `mc_rip` on FreeBSD/x86_64), but AArch64 resumes from
a signal at ELR rather than LR, and `gp_elr` was left holding the
interrupted pc.

Together these meant `powermod(1, 0, big(0))`, which reaches
`__gmp_exception` -> `raise(SIGFPE)`, returned from the handler back
into `__sys_thr_kill+8` with `x30` clobbered to 0, fell through to the
`ret` two instructions later, and branched to address 0:

    Fault at memory address: 0x0

    [6995] signal 11 (1): Segmentation fault
    in expression starting at none:1
    unknown function (ip: 0x0) at (unknown file)

While here, restore `d8`-`d15` into `q8`-`q15` rather than `q7`-`q14`,
which also dropped `d15`.

---

This commit, both its message and diff, were generated by AI. I, a
human, edited the message for formatting and have reviewed the overall
output and think it seems reasonable, though it falls a bit outside of
my expertise, so I'd appreciate additional human eyeballs.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…links

Fix dead JuliaHub blog links in the inference devdocs
…aLang#62993)

This is a noticeable improvement to nested reductions in the presence of
arrays.

It is a bit silly to have to work around the recursion heuristic in this
way, but I am not sure we can do much better unless we teach it to
associate an argument's type with the dispatch that is keyed on it
(these `mapreduce` methods are non-recursive since they immediately
invoke different Methods, but the compiler has a very hard time seeing
that from the types alone).

This makes many pithy reduce-heavy examples like `maximum(count(!iszero,
v .* i) for i in 1:3)` fully infer and `--trim`-compatible.

Prepared with assistance from Claude Fable 5 🤖

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…invokelatest` (JuliaLang#62416)

This brings `Core.invoke_in_world` to parity with `Core.invokelatest` by
eliminating the boxing of `world`:
```julia
julia> @Btime Core.invoke_in_world(Base.get_world_counter(), $(Base.setindex!), $(Ref{Any}()), nothing)
  9.952 ns (0 allocations: 0 bytes) # PR
  20.316 ns (1 allocation: 16 bytes) # master
Base.RefValue{Any}(nothing)

julia> @Btime Core.invokelatest($(Base.setindex!), $(Ref{Any}()), nothing)
  9.989 ns (0 allocations: 0 bytes) # PR
  13.960 ns (0 allocations: 0 bytes) # master
Base.RefValue{Any}(nothing)
```

a more typical invocation with lingering boxes for arguments / return
values:
```julia
julia> @Btime Core.invoke_in_world($w, $sin, $(Ref(1.0))[])
  17.201 ns (2 allocations: 32 bytes) # PR
  26.366 ns (3 allocations: 48 bytes) # master
```

Written by Claude Fable 5 🤖

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Brings in the SparseArrays changes that make its constructors, reductions,
and SuiteSparse solvers compile under `--trim`: JuliaSparse/SparseArrays.jl#751
(`sparse(I, J, V)` and `spdiagm` statically dispatchable),
JuliaSparse/SparseArrays.jl#753 (`mapreduce` and `sparsevec` specialized on
their function arguments), JuliaSparse/SparseArrays.jl#754 (`init_suitesparse`
trim-compatible), JuliaSparse/SparseArrays.jl#755 (CHOLMOD `RefValue` fields
and assertions), and JuliaSparse/SparseArrays.jl#756 (SPQR factors built with
typed constructors, and single precision input converted to double).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`jl_init_with_image_handle` receives an already-loaded image and sets
`jl_options.image_file` from `jl_pathname_for_handle`. It then called
`jl_resolve_sysimg_location(JL_IMAGE_JULIA_HOME, ...)`, which would
prefix a non-absolute `image_file` with `julia_bindir`. That is not the
right interpretation for an in-memory image; `JL_IMAGE_IN_MEMORY` exists
for this case and skips the re-rooting.

In practice the path returned by `jl_pathname_for_handle` is absolute,
so this is mostly a correctness/intent fix, but it matters for the
static-linking case where the "image" is the executable itself.

Part of the static libjulia-internal work (JuliaLang#62868 follow-up).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…porting" (JuliaLang#63014)

Reverts JuliaLang#62396

> Any impact on precompile time etc? Not to block, just to attribute if
so.

> Seems unlikely, since this generally only makes queries cheaper, and
neutral for small queries, and only matters for queries that land
somewhere difficult (multiple ambiguous matches). But best practice is
indeed to at least check if nanosoldier moves either direction:

This was not true
(JuliaLang#62396 (comment)).
Reverting based on the PR doing things it wasn't expect it to do.
Extend the `Trimmability` trim test with the SparseArrays functionality that
recent SparseArrays work made statically dispatchable: the `sparse(I, J, V)`,
`sparse` with `combine`, and `spdiagm` constructors; products, structural
operations, broadcasting and concatenation; sparse vectors; and reductions,
including the nested reductions that JuliaLang#62993 keeps type-stable. Each group
prints one line that the test compares against the untrimmed result.

The SuiteSparse-backed solves and factorizations are written but left
commented out: loading their libraries inside a trimmed binary still fails
until JuliaLang#62912 lands.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
JuliaC's trim buildscript overrides `_mapreduce_dim` for the closed world, and
until JuliaLang/JuliaC.jl#184 that override still routed array reductions with
an `init` through `mapfoldl_impl`, undoing the inference fix from JuliaLang#62993 in
trimmed images. Pin past it so the trim tests run against the same reduction
code that the sysimage uses.

Only the manifest moves, since the project already tracks `main`: this records
the resolved tree, JuliaC 0.3.10, which also replaces Patchelf_jll with
LIEF_Patchelf_jll and refreshes the recorded stdlib versions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…en-sizes

codegen: Apply odd-bit storage widths consistently
We were maintaining the allocated bytes in `gc_num.allocd` but not the
allocation count in `gc_num.poolalloc`. This should fix allocation
counts reported by `@allocations` and `@time` / `@timed`

These counts were always broken, but
JuliaLang#62905 introduced some codegen
changes that affected allocation elision that made the `misc.jl` tests
hit this by accident.

Found after investigation of
https://buildkite.com/julialang/julia-pr/builds/2070#01a06c76-9fbf-46a0-82cb-61c875d1e955/L1604
by Claude Opus 5 🤖
Apply the XLEN extension rule to 32-bit arguments and integer returns, including unsigned values such as PCRE2_ANCHORED. Flatten eligible aggregates for the hardware floating-point convention, excluding unions, pointers, boxed fields, and vectors, and derive ABI_FLEN from the runtime ABI.

Track the variadic tail so RISC-V uses the integer convention. Preserve 16-bit floating-point LLVM types for backend lowering, and restrict C default floating-point promotion to float.

Assisted-by: Codex (GPT-6)
Keep opaque 128-bit scalars as i128 so LLVM spills the whole argument when fewer than two integer registers remain. Classify 8-byte vectors as SSE while retaining the GCC memory convention for one-element floating-point vectors.

Assisted-by: Codex (GPT-6)
Recognize short vector tuples by their 8- or 16-byte size before unwrapping single-field aggregates, and exclude smaller vectors from HFA classification. Ordinary structs containing VecElement fields remain scalar aggregates. Check the current field type before recursing into an aggregate.

Keep 16-byte-aligned composites as i128 so LLVM preserves the target register and stack alignment rules, including Darwin varargs.

Assisted-by: Codex (GPT-6)
Exercise C interoperability in both directions for scalar, aggregate, vector, and variadic arguments and returns. Trailing sentinels detect incorrect register accounting as well as corrupted values.

Skip incompatible compiler vector conventions and x86-64 SysV register-save-area probes affected by the old GCC variadic __int128 bug. The latter caused the ccall failure in https://buildkite.com/julialang/julia-pr/builds/2021#01a066d7-9686-4295-81fe-e5b3f3058354.

Assisted-by: Codex (GPT-6)
Document the GCC convention used for one-element floating-point vectors on x86-64 and the distinction between short vectors and smaller composites on AArch64.

Assisted-by: Codex (GPT-6)
…ys-trim

deps: Bump `SparseArrays` to `67ff5ff` + add `--trim` regression tests
…liaLang#62944)

Debugged to this based on weird errors when trying to use julia with
`USE_TRACY` enabled:

The object cache is shared by builds using the same depot. `USE_TRACY`
changes `jl_task_t`, so a Tracy build can otherwise load cached code
with
offsets from a non-Tracy build and crash.

Fix this by hashing preprocessed sources.

Assisted-by: Codex (GPT-5)

---------

Co-authored-by: Valentin Churavy <v.churavy@gmail.com>
Co-authored-by: Sam Schweigel <me@xal.systems>
Reorder smalltags to speed up GC marking again
When smalltags were first added in e26b63c, kinds were made to
occupy a contiguous range, as well as tags for types that take the
generic marking path in the GC.  Since then, new tags have been added
that must also go through the generic GC path, which caused the code
generated for GC marking to regress.

The important case was the big condition at the beginning of
`gc_mark_outrefs`, checking for all of the smalltags that must be
scanned by the GC but which have no special handling:

```c
if (vtag == (jl_datatype_tag << 4) ||
    vtag == (jl_unionall_tag << 4) ||
    vtag == (jl_uniontype_tag << 4) ||
    vtag == (jl_typeeq_tag << 4) ||
    vtag == (jl_typeegal_tag << 4) ||
    vtag == (jl_tvar_tag << 4) ||
    vtag == (jl_vararg_tag << 4) ||
    vtag == (jl_globalref_tag << 4) ||
    vtag == (jl_gotoifnot_tag << 4) ||
    vtag == (jl_returnnode_tag << 4) ||
    vtag == (jl_enternode_tag << 4) ||
    vtag == (jl_pinode_tag << 4) ||
    vtag == (jl_phinode_tag << 4) ||
    vtag == (jl_phicnode_tag << 4) ||
    vtag == (jl_upsilonnode_tag << 4) ||
    vtag == (jl_quotenode_tag << 4)) {
    // these objects have pointers in them, but no other special handling
    // so we want these to fall through to the end
    vtag = (uintptr_t)ijl_small_typeof[vtag / sizeof(*ijl_small_typeof)];
}
```

The code generated for this was quite bad.  If the objects being scanned
are close enough to be cache hits, this made GC marking instruction
bound.  This commit puts the smalltags back into a sensible order, and
changes the check to a range check:

```c
if (vtag - (jl_gc_generic_tags_first << 4) <=
    (jl_gc_generic_tags_last - jl_gc_generic_tags_first) << 4)
```

On this microbenchmark, minimum GC times are improved by about 10%:
```julia
mutable struct Node
    next::Union{Nothing,Node}
    value::Int
end

function makeheap(n)
    nodes = [Node(nothing, i) for i in 1:n]
    order = collect(1:n)
    for i in 1:n-1
        nodes[i].next = nodes[i+1]
    end
    nodes[1]
end

GC.gc()
heap = makeheap(10_000_000)
times = Float64[]
for i=1:50
    push!(times, (@timed GC.gc()).gctime)
end

println(minimum(times))
```

Assisted-by: Codex (GPT-6)
This fixes JuliaLang#62296.

## MWE 

```julia
using Test

ftrue(; kw...) = true
ffalse(; kw...) = false
kw = (; a = 1)

@test ftrue(; kw...)
@test ffalse(; kw...)
```

## Before this PR

```julia
julia> @test ftrue(; kw...) # Same thing happens for `ffalse`.
Error During Test at REPL[4]:1
  Test threw exception
  Expression: ftrue(; kw...)
  BoundsError: attempt to access Int64 at index [2]
  Stacktrace:
   [1] indexed_iterate(I::Int64, i::Int64, state::Nothing)
     @ Base tuple.jl:170
   [2] merge(a::@NamedTuple{}, itr::Tuple{Int64})
     @ Base namedtuple.jl:368
   [3] eval_test_function(func::Any, args::Any, kwargs::Any, quoted_func::Union{Expr, Symbol}, source::LineNumberNode, negate::Bool)
     @ Test ~/jl/julia/stdlib/Test/src/Test.jl:407
   [4] top-level scope
     @ REPL[5]:1
   [5] macro expansion
     @ ~/jl/julia/stdlib/Test/src/Test.jl:781 [inlined]
```

## After this PR

```julia
julia> @test ftrue(; kw...)
Test Passed

julia> @test ffalse(; kw...)
Test Failed at REPL[5]:1
  Expression: ffalse(; kw...)
   Evaluated: ffalse(; a = 1)
ERROR: There was an error during testing
```
…iaLang#63036)

JuliaLICM can sink an inner loop's `gc_preserve_end` past the enclosing
loop containing its `gc_preserve_begin`. The token then escapes its
defining
loop, and LCSSA cannot insert token PHIs. Subsequent loop cloning can
break
dominance, as seen while precompiling Parsers 3.0.0 on Julia 1.12 with
LLVM
assertions in [JuliaGPU/GPUCompiler.jl#924's
CI](https://github.com/JuliaGPU/GPUCompiler.jl/actions/runs/34050391805/job/101532856212).

Restrict sinking to exits inside the begin's loop. Begins outside all
loops
remain unrestricted. Omitting an end conservatively extends
preservation,
which GC lowering already supports.

Assisted-by: Claude Fable 5.1 and GPT 6 Astra
topolarity and others added 28 commits September 22, 2026 12:14
This includes a few changes to our write barrier APIs:
  - Remove `jl_gc_notify_task_suspend` which was a vestigial post-write
    barrier to make up for `jl_gc_wb_current_task` being disabled on
    MMTk StickyImmix. Barrier on the task consistently instead.
  - Add `jl_gc_wb_object` as a whole-object write barrier, which does
    not inspect the age etc. of the new pointee. On field-precise
    collectors this marks all object fields as dirty.
  - Remove `jl_gc_wb_back` which was similar to a whole-object barrier
    but defined with post-write semantics, where it would inspect the
    fields of the object to look for young pointees. That whole-object
    scan is better handled / amortized inside the GC anyway.
  - Move `jl_gc_wb_cold` and `jl_gc_queue_multiroot` to be internal-only
    APIs (move them out of gc-interface.h) and rename
    `jl_gc_queue_multiroot` to `jl_gc_multi_wb_cold` for consistency

Assisted-by: Claude Code (Fable 5)
…parseArrays-f1177ba-master

🤖 Bump SparseArrays stdlib 3e67d6e → f1177ba
Julia's profiler on Windows runs a dedicated sampling thread. Every
millisecond it visits each Julia thread in turn and, for one thread at a
time, calls `SuspendThread`, reads the register context, unwinds that
thread's stack, and resumes it. Unwinding calls into the Windows
runtime, so the sampler deadlocks whenever the thread it froze holds a
lock the unwinder then needs - `RtlAllocateHeap` and `LdrLoadDll` are
the ones in the reports. JuliaLang#59877 added the only practical escape: a
one-shot timer that fires 1s into a stuck unwind, resumes the frozen
thread and makes the unwinder abandon that sample. JuliaLang#60463 moved it to
the timer-queue API after the original registration turned out to drop
callbacks under contention.

All of it is `#ifdef _CPU_X86_64_`. On i686 the sampler calls dbghelp's
`StackWalk64` with the target suspended and has no way out, and dbghelp
allocates from the process heap - the same `RtlAllocateHeap` lock the
x86-64 reports name. This is the leading suspect for the occasional 200s
timeout of the "spawn and wait lots of tasks" test on i686-w64-mingw32,
which profiles two million task spawns; that has not been confirmed on a
hung machine.

Arm the watchdog on both architectures and give the i686 `StackWalk64`
calls the same abort window, now shared by both unwinders rather than
open-coded.

Two fixes the shared path needs. `jl_set_profile_abort_ptr` touches an
emulated-TLS variable, and `__emutls_get_address` calls `calloc` on
first use - the escape hatch taking the very heap lock it exists to
escape, which is where the trace in JuliaLang#60454 hangs; allocate the block
when the sampler thread starts instead. A failed suspend left the timer
armed, so it fired against a later sample's window and resumed a thread
nobody had suspended.

Assisted-by: Claude Code (Opus 5)
contrib: add optimized Windows builds and fix Linux/macOS regressions
- `Meta.parse` and friends should parse using VERSION by default

- Be slightly stricter when flisp parsing is asked for non-VERSION syntax

- `__toplevel__.var"#_internal_julia_parse"` should be probably be the versioned
  kind, though I can't tell if this causes issues today

- Move JuliaSyntax activation to before stdlib precompilation, and check on
    startup to disable it rather than enable it.  Note that I don't check
    JULIA_USE_FLISP_PARSER while building, as the flisp parser doesn't support
    the 1.13 version that at least one stdlib asks for.

- Remove the debug logging kwarg from `JuliaSyntax.enable_in_core!`---I don't
    know if any of this machinery is needed anymore, so I've kept it around, but
    the entrypoint in Base.jl was messing with trimmability

Some tests and fixes by claude

Assisted-by: claude fable 5.1
Stage 0 installs BOLT from BOLT_jll like the rest of the toolchain, but
the stage-0 comment and the README still describe a source build.
BOLT rejects -split-strategy=cdsplit with AArch64's default code model,
and -jump-tables=move has no effect because BOLT does not reconstruct
AArch64 jump tables. Limit these options to x86-64, preserving its
existing flags and the libjulia-internal splitting exception.
On AArch64, BOLT could not instrument or rewrite a ThinLTO-built libLLVM:
InstCombinePass::run grows past 1 MB, and BOLT then tried to relax every
ADR in it, which it cannot do in a function whose jump tables it does not
recover ("cannot relax ADR in non-simple function"). BOLT_jll v23.1.2+0
carries the upstream fix (llvm/llvm-project#215415), which is not in any
release yet, so libLLVM can now be rewritten on AArch64 as on x86-64.

Record checksums for every tarball that BOLT_jll v23.1.2+0 ships: x86_64
and aarch64, glibc and musl. Only the x86_64 glibc tarball had one before,
and release builds refuse to autogenerate a missing checksum. Source
builds (USE_BINARYBUILDER_BOLT=0) apply the same patch.
…g#63311)

`reserve_module_binding_i` scanned `"$basename##i"` from `i = 0` on
every call, which is quadratic for a long-lived process that keeps
lowering into the same module (e.g. JETLS on Julia 1.13, where
`module_unique_name` falls back to it with the shared basename `""`).
Probe `0, 1, 3, 7, ...` for a free index and bisect back to the
smallest one instead, so each call takes $O(log n)$ probes while the
generated names stay unchanged for contiguous reservations.
Support Linux AArch64 in the unified optimized build
aotcompile: shard 32-bit image builds, compiling one shard at a time
Without BinaryBuilder's CSL, the GCC runtime libraries are copied from the
compiler's library search path, and so was libpthread. On Linux that is
part of glibc, not of GCC: with a toolchain that has its own sysroot (e.g.
an old glibc, as the CI images' toolchain), that copy is incompatible
with the system's libc, and every program with a usr/lib rpath fails to
run (e.g. curl's configure: "cannot run C compiled programs").
LLVM's tools link against libLLVM, which depends on our zlib and zstd.
GNU ld looks for the dependencies of shared libraries in -rpath-link
directories and its default ones, not in -L directories, so with a
toolchain whose default search path is its own sysroot, linking fails
with undefined references to zlib/zstd functions.
…ing is stale (JuliaLang#63317)

* loading: record a cache file's own freshness verdict in `stale_cache`

`compilecache_freshest_path` memoizes the file-level verdicts of a
package's dependencies in `stale_cache`, but not the verdict for the
package's own cache file. In the precompilation driver every package is
checked after its dependencies, so the first dependent of each package
re-ran `stale_cachefile` on a file that had just been validated,
re-parsing its header and re-checksumming both the `.ji` and the
pkgimage. On a no-op run over a 297-package environment this doubled the
checksum work for 273 of the 297 cache files.

Record the verdict under the build id `stale_cachefile` returns, which
is the key a dependent looks it up by.

Assisted-by: Claude Code (Opus 5.5)

* precompilation: compute scheduling priorities on demand

The critical-path priorities from JuliaLang#63016 were computed up front for
every package in the environment, and the cost estimate behind them
walks each package's source tree. That ran on every call, including
no-op ones where nothing needs a worker, and cost about 40 ms on a
297-package environment (15-25% of a no-op `precompilepkgs`).

Compute a package's priority when it asks for a worker slot instead.
Only stale packages ask, and the height recursion only visits their
dependents, which are almost always stale too. The priorities are
unchanged.

Assisted-by: Claude Code (Opus 5.5)
…3324)

`typed_load` loads an atomic field whose type is an inline immutable
with GC pointers as an integer into the `atomic_load_box` alloca. The
value was reloaded with its pointer-exposing type only when the field
might be null; otherwise the alloca itself was returned as a slot,
whose pointers, being stored as an integer, are not rooted. When the
slot is address-taken, e.g. passed by reference to a non-inlined call,
nothing keeps those pointers alive, so overwriting the field and
running a GC could free objects still reachable from the loaded value.

Always reload the value with its pointer-exposing type when it
contains GC pointers, as `typed_store` already does for the old value
it returns.

Fixes JuliaLang#63320

Assisted-by: Claude Code (Opus 5.5)
macOS's ld64 rejects it ("ld: unknown options: -rpath-link"), which broke
LLVM source builds there; scope it like Make.inc's RPATH.
* Accelerate mixed-signedness `BitInteger` comparisons

As found by [@matthias314](JuliaLang#63287 (comment))
`<`, `<=`, `>` and `>=` can be accelerated for `BitSigned`-`BigUnsigned` combinations of different widths.

:robot: :
Compare in the signed type when it can represent every value of the unsigned type, avoiding redundant sign checks. Share comparison logic across operand orders.

Benchmarks on an Intel Core i7-1270P, pinned to a performance core, using Julia 1.14.0-DEV.3129 with LLVM 22.1.8. Both tables measure warmed 256-element arrays with Bool output. Scalar measurements disable LLVM loop and SLP vectorization. Timings exclude input generation but include loads, stores, and loop overhead.

Values are changes in execution time, rounded to whole percentages. Negative means faster. `S` denotes the signed operand and `U` the unsigned operand.

### Scalar Loops

| Operand types | S < U | U < S | S <= U | U <= S |
|---|---:|---:|---:|---:|
| Int16 / UInt8 | -44% | -44% | -42% | -45% |
| Int32 / UInt8 | -45% | -44% | -44% | -44% |
| Int32 / UInt16 | -45% | -43% | -43% | -43% |
| Int64 / UInt8 | -43% | -43% | -41% | -43% |
| Int64 / UInt16 | -43% | -43% | -42% | -44% |
| Int64 / UInt32 | -42% | -42% | -41% | -41% |
| Int128 / UInt8 | -32% | -34% | -34% | -25% |
| Int128 / UInt16 | -27% | -32% | -35% | -28% |
| Int128 / UInt32 | -26% | -33% | -35% | -25% |
| Int128 / UInt64 | -27% | -35% | -38% | -25% |

### Vector Loops

| Operand types | S < U | U < S | S <= U | U <= S |
|---|---:|---:|---:|---:|
| Int16 / UInt8 | -10% | -2% | -2% | 0% |
| Int32 / UInt8 | 0% | +10% | +10% | +10% |
| Int32 / UInt16 | 0% | +10% | +10% | +10% |
| Int64 / UInt8 | -19% | -19% | -15% | -15% |
| Int64 / UInt16 | -19% | -19% | -15% | -15% |
| Int64 / UInt32 | -19% | -19% | -16% | -15% |
| Int128 / UInt8 | -46% | -40% | -46% | -39% |
| Int128 / UInt16 | -46% | -39% | -46% | -40% |
| Int128 / UInt32 | -48% | -41% | -48% | -41% |
| Int128 / UInt64 | -41% | -45% | -49% | -34% |

Measurements compare averages of the fastest half of 3,001 samples, with 1,000 passes per sample and both benchmark orders checked. Identical-function controls were below 1% for all retained vector cases and 32 of 40 scalar cases. The remaining eight scalar controls differed by roughly 1–6%, so the percentages should be interpreted as approximate.

The Int32 vector regressions result from LLVM choosing less efficient Boolean packing. Changing only that packing sequence removes the regression in an isolated assembly experiment. Reverting the operand-forwarding simplification produces identical vector assembly and therefore does not help.

Overall, these workloads favor the optimization, although the benefit depends on operand types and execution mode. Ultimately, target-aware comparison optimization and Boolean packing belong in LLVM.

* Make type conversion explicit

Thanks for the suggestion, @vtjnash!

---------

Co-authored-by: Patrick Häcker <patrick.haecker@bosch.com>
* REPL: Sync the ^C sweep test on the evaluation result

In julia-pr build 2541 this test hung at the `Cancelled all in-flight
work.` read and was hard-killed after the 900s timeout, on both
i686-linux-gnu and aarch64-linux-gnu:
https://buildkite.com/julialang/julia-pr/builds/2541#01a0aee0-fc2e-43c8-a079-4161fe774fa9
The sweep had in fact thrown - a `CancellationTokenSource` walk reached
freed memory, since fixed by JuliaLang#63234 - and the keymap logs and swallows
that, so the message never came and the read blocked forever.

The test widened that window. Its stand-down step types `1` and reads
until `julia> `, but the `^C` handler that armed the sweep ends in
`transition(s, :reset); refresh_line(s)`, so a prompt is already sitting
unread in the pipe and the read matches that instead, before `1` has
been processed at all. The step meant to stand the arm down is never
waited for, and everything after it races the evaluation. Looping the
block under `-t1` hit the freed-memory error 5 times in 600 iterations
as written, and once in 900 with the read below.

Communicate through a result value, as the rest of the block does; the
concatenation keeps the needle out of LineEdit's echo of the typed
characters.

Assisted-by: Claude Code (Opus 5)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* REPL: Reword the double-^C test comments

Review of this test found its comments hard to read: they lean on
invented vocabulary ("arms the sweep", "stands the arm down", "session
epoch") instead of saying what each key press does. Describe the
behaviour in plain terms, and spell out why the awaited results are
written as concatenations. The marker expression is renamed to match.

Assisted-by: Claude Code (Fable 5.1)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…uild

build: fix USE_BINARYBUILDER=0 builds with sysroot toolchains
@qinsoon

qinsoon commented Sep 28, 2026

Copy link
Copy Markdown
Member

Can you reopen the PR? Target JuliaLang/Julia (not mmtk/julia). We should consider archive our fork to avoid this from happening again.

@wks wks closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.