Repository navigation
Conversation
…62998) This reverts PR JuliaLang#62806 / commit 294c405. Alternative to and hence closes JuliaLang#62996
`jl_init_options` is normally called by the libjulia loader (`cli/loader_lib.c`) before any other runtime code. In a static build of libjulia-internal (JuliaLang#62868) there is no loader, and `jl_autoinit_and_adopt_thread` — the trampoline that images call on first entry — becomes the first runtime entry point, so `jl_options` would never be initialized. Call `jl_init_options` there before `jl_init_with_image_handle`. `jl_init_options` guards itself with `jl_options_initialized`, so this is a no-op for the regular shared build. Part of the static libjulia-internal work. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Emit load/store pairs for small padding-free pointer-free values, keeping source and destination alias metadata separate. Helps quite a bit with vectorization over a variety of common types ```julia Base.@noinline function _copyto_bench!(dst, src) for i in eachindex(dst, src) dst[i] = src[i] end dst end n = 1024 T = Tuple{Bool, UInt8} src = (rand(T, n)); dst = copy(src); @Btime _copyto_bench!($dst, $src); ``` <img width="866" height="323" alt="image" src="https://github.com/user-attachments/assets/ff96c25a-d443-4163-8874-bb23688c9ca3" /> Fixes JuliaLang#60409 Coauthored-by: Codex (GPT-5.6 Sol xhigh)
…candidates (JuliaLang#62396) Inspired by JuliaLang#62262, Claude Fable identified a bug in the current algorithm, as well as proposed some improvements for the accuracy of the ambiguity computation. This still needs some human intervention to make sure it does things correctly / efficiently (note that one of the commits is just NFC code moving). --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`jl_simulate_longjmp` resolves an in-signal throw by decoding the target
`jl_jmp_buf` into the interrupted `ucontext` and letting `sigreturn`
perform the jump. The FreeBSD/AArch64 implementation got this wrong in
two ways, so every throw out of a signal handler—`DivideError`,
`StackOverflowError`, `ReadOnlyMemoryError`, and the `safe_restore`
recovery used by `jl_` and the profiler—crashed the process instead.
First, the jump buffer offsets were off by two. FreeBSD declares the
buffer as `__int128_t _sjb[]` in `<machine/setjmp.h>`, but that element
type only forces the size and alignment; `_setjmp` stores packed 8-byte
words:
[0] magic, [1] sp, [2..13] x19..x30, [14..21] d8..d15
Decoding from index 0 onward put the `setjmp` magic into `x19`, the
saved `x28` into `gp_lr`, and the saved frame pointer into `gp_sp`. The
`assert(gp_sp % 16 == 0)` sanity check passed regardless, since a frame
pointer is 16-byte aligned too.
Second, the resume pc was never set. Every other architecture ends by
assigning the restored return address to the pc (`mc->pc = mc->regs[30]`
on Linux/AArch64, `mc_rip` on FreeBSD/x86_64), but AArch64 resumes from
a signal at ELR rather than LR, and `gp_elr` was left holding the
interrupted pc.
Together these meant `powermod(1, 0, big(0))`, which reaches
`__gmp_exception` -> `raise(SIGFPE)`, returned from the handler back
into `__sys_thr_kill+8` with `x30` clobbered to 0, fell through to the
`ret` two instructions later, and branched to address 0:
Fault at memory address: 0x0
[6995] signal 11 (1): Segmentation fault
in expression starting at none:1
unknown function (ip: 0x0) at (unknown file)
While here, restore `d8`-`d15` into `q8`-`q15` rather than `q7`-`q14`,
which also dropped `d15`.
---
This commit, both its message and diff, were generated by AI. I, a
human, edited the message for formatting and have reviewed the overall
output and think it seems reasonable, though it falls a bit outside of
my expertise, so I'd appreciate additional human eyeballs.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…links Fix dead JuliaHub blog links in the inference devdocs
…aLang#62993) This is a noticeable improvement to nested reductions in the presence of arrays. It is a bit silly to have to work around the recursion heuristic in this way, but I am not sure we can do much better unless we teach it to associate an argument's type with the dispatch that is keyed on it (these `mapreduce` methods are non-recursive since they immediately invoke different Methods, but the compiler has a very hard time seeing that from the types alone). This makes many pithy reduce-heavy examples like `maximum(count(!iszero, v .* i) for i in 1:3)` fully infer and `--trim`-compatible. Prepared with assistance from Claude Fable 5 🤖 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…invokelatest` (JuliaLang#62416) This brings `Core.invoke_in_world` to parity with `Core.invokelatest` by eliminating the boxing of `world`: ```julia julia> @Btime Core.invoke_in_world(Base.get_world_counter(), $(Base.setindex!), $(Ref{Any}()), nothing) 9.952 ns (0 allocations: 0 bytes) # PR 20.316 ns (1 allocation: 16 bytes) # master Base.RefValue{Any}(nothing) julia> @Btime Core.invokelatest($(Base.setindex!), $(Ref{Any}()), nothing) 9.989 ns (0 allocations: 0 bytes) # PR 13.960 ns (0 allocations: 0 bytes) # master Base.RefValue{Any}(nothing) ``` a more typical invocation with lingering boxes for arguments / return values: ```julia julia> @Btime Core.invoke_in_world($w, $sin, $(Ref(1.0))[]) 17.201 ns (2 allocations: 32 bytes) # PR 26.366 ns (3 allocations: 48 bytes) # master ``` Written by Claude Fable 5 🤖 --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Brings in the SparseArrays changes that make its constructors, reductions, and SuiteSparse solvers compile under `--trim`: JuliaSparse/SparseArrays.jl#751 (`sparse(I, J, V)` and `spdiagm` statically dispatchable), JuliaSparse/SparseArrays.jl#753 (`mapreduce` and `sparsevec` specialized on their function arguments), JuliaSparse/SparseArrays.jl#754 (`init_suitesparse` trim-compatible), JuliaSparse/SparseArrays.jl#755 (CHOLMOD `RefValue` fields and assertions), and JuliaSparse/SparseArrays.jl#756 (SPQR factors built with typed constructors, and single precision input converted to double). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`jl_init_with_image_handle` receives an already-loaded image and sets `jl_options.image_file` from `jl_pathname_for_handle`. It then called `jl_resolve_sysimg_location(JL_IMAGE_JULIA_HOME, ...)`, which would prefix a non-absolute `image_file` with `julia_bindir`. That is not the right interpretation for an in-memory image; `JL_IMAGE_IN_MEMORY` exists for this case and skips the re-rooting. In practice the path returned by `jl_pathname_for_handle` is absolute, so this is mostly a correctness/intent fix, but it matters for the static-linking case where the "image" is the executable itself. Part of the static libjulia-internal work (JuliaLang#62868 follow-up). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…porting" (JuliaLang#63014) Reverts JuliaLang#62396 > Any impact on precompile time etc? Not to block, just to attribute if so. > Seems unlikely, since this generally only makes queries cheaper, and neutral for small queries, and only matters for queries that land somewhere difficult (multiple ambiguous matches). But best practice is indeed to at least check if nanosoldier moves either direction: This was not true (JuliaLang#62396 (comment)). Reverting based on the PR doing things it wasn't expect it to do.
Extend the `Trimmability` trim test with the SparseArrays functionality that recent SparseArrays work made statically dispatchable: the `sparse(I, J, V)`, `sparse` with `combine`, and `spdiagm` constructors; products, structural operations, broadcasting and concatenation; sparse vectors; and reductions, including the nested reductions that JuliaLang#62993 keeps type-stable. Each group prints one line that the test compares against the untrimmed result. The SuiteSparse-backed solves and factorizations are written but left commented out: loading their libraries inside a trimmed binary still fails until JuliaLang#62912 lands. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
JuliaC's trim buildscript overrides `_mapreduce_dim` for the closed world, and until JuliaLang/JuliaC.jl#184 that override still routed array reductions with an `init` through `mapfoldl_impl`, undoing the inference fix from JuliaLang#62993 in trimmed images. Pin past it so the trim tests run against the same reduction code that the sysimage uses. Only the manifest moves, since the project already tracks `main`: this records the resolved tree, JuliaC 0.3.10, which also replaces Patchelf_jll with LIEF_Patchelf_jll and refreshes the recorded stdlib versions. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…en-sizes codegen: Apply odd-bit storage widths consistently
We were maintaining the allocated bytes in `gc_num.allocd` but not the allocation count in `gc_num.poolalloc`. This should fix allocation counts reported by `@allocations` and `@time` / `@timed` These counts were always broken, but JuliaLang#62905 introduced some codegen changes that affected allocation elision that made the `misc.jl` tests hit this by accident. Found after investigation of https://buildkite.com/julialang/julia-pr/builds/2070#01a06c76-9fbf-46a0-82cb-61c875d1e955/L1604 by Claude Opus 5 🤖
Apply the XLEN extension rule to 32-bit arguments and integer returns, including unsigned values such as PCRE2_ANCHORED. Flatten eligible aggregates for the hardware floating-point convention, excluding unions, pointers, boxed fields, and vectors, and derive ABI_FLEN from the runtime ABI. Track the variadic tail so RISC-V uses the integer convention. Preserve 16-bit floating-point LLVM types for backend lowering, and restrict C default floating-point promotion to float. Assisted-by: Codex (GPT-6)
Keep opaque 128-bit scalars as i128 so LLVM spills the whole argument when fewer than two integer registers remain. Classify 8-byte vectors as SSE while retaining the GCC memory convention for one-element floating-point vectors. Assisted-by: Codex (GPT-6)
Recognize short vector tuples by their 8- or 16-byte size before unwrapping single-field aggregates, and exclude smaller vectors from HFA classification. Ordinary structs containing VecElement fields remain scalar aggregates. Check the current field type before recursing into an aggregate. Keep 16-byte-aligned composites as i128 so LLVM preserves the target register and stack alignment rules, including Darwin varargs. Assisted-by: Codex (GPT-6)
Exercise C interoperability in both directions for scalar, aggregate, vector, and variadic arguments and returns. Trailing sentinels detect incorrect register accounting as well as corrupted values. Skip incompatible compiler vector conventions and x86-64 SysV register-save-area probes affected by the old GCC variadic __int128 bug. The latter caused the ccall failure in https://buildkite.com/julialang/julia-pr/builds/2021#01a066d7-9686-4295-81fe-e5b3f3058354. Assisted-by: Codex (GPT-6)
Document the GCC convention used for one-element floating-point vectors on x86-64 and the distinction between short vectors and smaller composites on AArch64. Assisted-by: Codex (GPT-6)
…ys-trim deps: Bump `SparseArrays` to `67ff5ff` + add `--trim` regression tests
…liaLang#62944) Debugged to this based on weird errors when trying to use julia with `USE_TRACY` enabled: The object cache is shared by builds using the same depot. `USE_TRACY` changes `jl_task_t`, so a Tracy build can otherwise load cached code with offsets from a non-Tracy build and crash. Fix this by hashing preprocessed sources. Assisted-by: Codex (GPT-5) --------- Co-authored-by: Valentin Churavy <v.churavy@gmail.com> Co-authored-by: Sam Schweigel <me@xal.systems>
Reorder smalltags to speed up GC marking again When smalltags were first added in e26b63c, kinds were made to occupy a contiguous range, as well as tags for types that take the generic marking path in the GC. Since then, new tags have been added that must also go through the generic GC path, which caused the code generated for GC marking to regress. The important case was the big condition at the beginning of `gc_mark_outrefs`, checking for all of the smalltags that must be scanned by the GC but which have no special handling: ```c if (vtag == (jl_datatype_tag << 4) || vtag == (jl_unionall_tag << 4) || vtag == (jl_uniontype_tag << 4) || vtag == (jl_typeeq_tag << 4) || vtag == (jl_typeegal_tag << 4) || vtag == (jl_tvar_tag << 4) || vtag == (jl_vararg_tag << 4) || vtag == (jl_globalref_tag << 4) || vtag == (jl_gotoifnot_tag << 4) || vtag == (jl_returnnode_tag << 4) || vtag == (jl_enternode_tag << 4) || vtag == (jl_pinode_tag << 4) || vtag == (jl_phinode_tag << 4) || vtag == (jl_phicnode_tag << 4) || vtag == (jl_upsilonnode_tag << 4) || vtag == (jl_quotenode_tag << 4)) { // these objects have pointers in them, but no other special handling // so we want these to fall through to the end vtag = (uintptr_t)ijl_small_typeof[vtag / sizeof(*ijl_small_typeof)]; } ``` The code generated for this was quite bad. If the objects being scanned are close enough to be cache hits, this made GC marking instruction bound. This commit puts the smalltags back into a sensible order, and changes the check to a range check: ```c if (vtag - (jl_gc_generic_tags_first << 4) <= (jl_gc_generic_tags_last - jl_gc_generic_tags_first) << 4) ``` On this microbenchmark, minimum GC times are improved by about 10%: ```julia mutable struct Node next::Union{Nothing,Node} value::Int end function makeheap(n) nodes = [Node(nothing, i) for i in 1:n] order = collect(1:n) for i in 1:n-1 nodes[i].next = nodes[i+1] end nodes[1] end GC.gc() heap = makeheap(10_000_000) times = Float64[] for i=1:50 push!(times, (@timed GC.gc()).gctime) end println(minimum(times)) ``` Assisted-by: Codex (GPT-6)
This fixes JuliaLang#62296. ## MWE ```julia using Test ftrue(; kw...) = true ffalse(; kw...) = false kw = (; a = 1) @test ftrue(; kw...) @test ffalse(; kw...) ``` ## Before this PR ```julia julia> @test ftrue(; kw...) # Same thing happens for `ffalse`. Error During Test at REPL[4]:1 Test threw exception Expression: ftrue(; kw...) BoundsError: attempt to access Int64 at index [2] Stacktrace: [1] indexed_iterate(I::Int64, i::Int64, state::Nothing) @ Base tuple.jl:170 [2] merge(a::@NamedTuple{}, itr::Tuple{Int64}) @ Base namedtuple.jl:368 [3] eval_test_function(func::Any, args::Any, kwargs::Any, quoted_func::Union{Expr, Symbol}, source::LineNumberNode, negate::Bool) @ Test ~/jl/julia/stdlib/Test/src/Test.jl:407 [4] top-level scope @ REPL[5]:1 [5] macro expansion @ ~/jl/julia/stdlib/Test/src/Test.jl:781 [inlined] ``` ## After this PR ```julia julia> @test ftrue(; kw...) Test Passed julia> @test ffalse(; kw...) Test Failed at REPL[5]:1 Expression: ffalse(; kw...) Evaluated: ffalse(; a = 1) ERROR: There was an error during testing ```
…iaLang#63036) JuliaLICM can sink an inner loop's `gc_preserve_end` past the enclosing loop containing its `gc_preserve_begin`. The token then escapes its defining loop, and LCSSA cannot insert token PHIs. Subsequent loop cloning can break dominance, as seen while precompiling Parsers 3.0.0 on Julia 1.12 with LLVM assertions in [JuliaGPU/GPUCompiler.jl#924's CI](https://github.com/JuliaGPU/GPUCompiler.jl/actions/runs/34050391805/job/101532856212). Restrict sinking to exits inside the begin's loop. Begins outside all loops remain unrestricted. Omitting an end conservatively extends preservation, which GC lowering already supports. Assisted-by: Claude Fable 5.1 and GPT 6 Astra
This includes a few changes to our write barrier APIs:
- Remove `jl_gc_notify_task_suspend` which was a vestigial post-write
barrier to make up for `jl_gc_wb_current_task` being disabled on
MMTk StickyImmix. Barrier on the task consistently instead.
- Add `jl_gc_wb_object` as a whole-object write barrier, which does
not inspect the age etc. of the new pointee. On field-precise
collectors this marks all object fields as dirty.
- Remove `jl_gc_wb_back` which was similar to a whole-object barrier
but defined with post-write semantics, where it would inspect the
fields of the object to look for young pointees. That whole-object
scan is better handled / amortized inside the GC anyway.
- Move `jl_gc_wb_cold` and `jl_gc_queue_multiroot` to be internal-only
APIs (move them out of gc-interface.h) and rename
`jl_gc_queue_multiroot` to `jl_gc_multi_wb_cold` for consistency
Assisted-by: Claude Code (Fable 5)
…parseArrays-f1177ba-master 🤖 Bump SparseArrays stdlib 3e67d6e → f1177ba
Julia's profiler on Windows runs a dedicated sampling thread. Every millisecond it visits each Julia thread in turn and, for one thread at a time, calls `SuspendThread`, reads the register context, unwinds that thread's stack, and resumes it. Unwinding calls into the Windows runtime, so the sampler deadlocks whenever the thread it froze holds a lock the unwinder then needs - `RtlAllocateHeap` and `LdrLoadDll` are the ones in the reports. JuliaLang#59877 added the only practical escape: a one-shot timer that fires 1s into a stuck unwind, resumes the frozen thread and makes the unwinder abandon that sample. JuliaLang#60463 moved it to the timer-queue API after the original registration turned out to drop callbacks under contention. All of it is `#ifdef _CPU_X86_64_`. On i686 the sampler calls dbghelp's `StackWalk64` with the target suspended and has no way out, and dbghelp allocates from the process heap - the same `RtlAllocateHeap` lock the x86-64 reports name. This is the leading suspect for the occasional 200s timeout of the "spawn and wait lots of tasks" test on i686-w64-mingw32, which profiles two million task spawns; that has not been confirmed on a hung machine. Arm the watchdog on both architectures and give the i686 `StackWalk64` calls the same abort window, now shared by both unwinders rather than open-coded. Two fixes the shared path needs. `jl_set_profile_abort_ptr` touches an emulated-TLS variable, and `__emutls_get_address` calls `calloc` on first use - the escape hatch taking the very heap lock it exists to escape, which is where the trace in JuliaLang#60454 hangs; allocate the block when the sampler thread starts instead. A failed suspend left the timer armed, so it fired against a later sample's window and resumed a thread nobody had suspended. Assisted-by: Claude Code (Opus 5)
contrib: add optimized Windows builds and fix Linux/macOS regressions
- `Meta.parse` and friends should parse using VERSION by default
- Be slightly stricter when flisp parsing is asked for non-VERSION syntax
- `__toplevel__.var"#_internal_julia_parse"` should be probably be the versioned
kind, though I can't tell if this causes issues today
- Move JuliaSyntax activation to before stdlib precompilation, and check on
startup to disable it rather than enable it. Note that I don't check
JULIA_USE_FLISP_PARSER while building, as the flisp parser doesn't support
the 1.13 version that at least one stdlib asks for.
- Remove the debug logging kwarg from `JuliaSyntax.enable_in_core!`---I don't
know if any of this machinery is needed anymore, so I've kept it around, but
the entrypoint in Base.jl was messing with trimmability
Some tests and fixes by claude
Assisted-by: claude fable 5.1
Stage 0 installs BOLT from BOLT_jll like the rest of the toolchain, but the stage-0 comment and the README still describe a source build.
BOLT rejects -split-strategy=cdsplit with AArch64's default code model, and -jump-tables=move has no effect because BOLT does not reconstruct AArch64 jump tables. Limit these options to x86-64, preserving its existing flags and the libjulia-internal splitting exception.
On AArch64, BOLT could not instrument or rewrite a ThinLTO-built libLLVM:
InstCombinePass::run grows past 1 MB, and BOLT then tried to relax every
ADR in it, which it cannot do in a function whose jump tables it does not
recover ("cannot relax ADR in non-simple function"). BOLT_jll v23.1.2+0
carries the upstream fix (llvm/llvm-project#215415), which is not in any
release yet, so libLLVM can now be rewritten on AArch64 as on x86-64.
Record checksums for every tarball that BOLT_jll v23.1.2+0 ships: x86_64
and aarch64, glibc and musl. Only the x86_64 glibc tarball had one before,
and release builds refuse to autogenerate a missing checksum. Source
builds (USE_BINARYBUILDER_BOLT=0) apply the same patch.
…g#63311) `reserve_module_binding_i` scanned `"$basename##i"` from `i = 0` on every call, which is quadratic for a long-lived process that keeps lowering into the same module (e.g. JETLS on Julia 1.13, where `module_unique_name` falls back to it with the shared basename `""`). Probe `0, 1, 3, 7, ...` for a free index and bisect back to the smallest one instead, so each call takes $O(log n)$ probes while the generated names stay unchanged for contiguous reservations.
Support Linux AArch64 in the unified optimized build
aotcompile: shard 32-bit image builds, compiling one shard at a time
Without BinaryBuilder's CSL, the GCC runtime libraries are copied from the compiler's library search path, and so was libpthread. On Linux that is part of glibc, not of GCC: with a toolchain that has its own sysroot (e.g. an old glibc, as the CI images' toolchain), that copy is incompatible with the system's libc, and every program with a usr/lib rpath fails to run (e.g. curl's configure: "cannot run C compiled programs").
LLVM's tools link against libLLVM, which depends on our zlib and zstd. GNU ld looks for the dependencies of shared libraries in -rpath-link directories and its default ones, not in -L directories, so with a toolchain whose default search path is its own sysroot, linking fails with undefined references to zlib/zstd functions.
…ing is stale (JuliaLang#63317) * loading: record a cache file's own freshness verdict in `stale_cache` `compilecache_freshest_path` memoizes the file-level verdicts of a package's dependencies in `stale_cache`, but not the verdict for the package's own cache file. In the precompilation driver every package is checked after its dependencies, so the first dependent of each package re-ran `stale_cachefile` on a file that had just been validated, re-parsing its header and re-checksumming both the `.ji` and the pkgimage. On a no-op run over a 297-package environment this doubled the checksum work for 273 of the 297 cache files. Record the verdict under the build id `stale_cachefile` returns, which is the key a dependent looks it up by. Assisted-by: Claude Code (Opus 5.5) * precompilation: compute scheduling priorities on demand The critical-path priorities from JuliaLang#63016 were computed up front for every package in the environment, and the cost estimate behind them walks each package's source tree. That ran on every call, including no-op ones where nothing needs a worker, and cost about 40 ms on a 297-package environment (15-25% of a no-op `precompilepkgs`). Compute a package's priority when it asks for a worker slot instead. Only stale packages ask, and the height recursion only visits their dependents, which are almost always stale too. The priorities are unchanged. Assisted-by: Claude Code (Opus 5.5)
…3324) `typed_load` loads an atomic field whose type is an inline immutable with GC pointers as an integer into the `atomic_load_box` alloca. The value was reloaded with its pointer-exposing type only when the field might be null; otherwise the alloca itself was returned as a slot, whose pointers, being stored as an integer, are not rooted. When the slot is address-taken, e.g. passed by reference to a non-inlined call, nothing keeps those pointers alive, so overwriting the field and running a GC could free objects still reachable from the loaded value. Always reload the value with its pointer-exposing type when it contains GC pointers, as `typed_store` already does for the old value it returns. Fixes JuliaLang#63320 Assisted-by: Claude Code (Opus 5.5)
macOS's ld64 rejects it ("ld: unknown options: -rpath-link"), which broke
LLVM source builds there; scope it like Make.inc's RPATH.
* Accelerate mixed-signedness `BitInteger` comparisons As found by [@matthias314](JuliaLang#63287 (comment)) `<`, `<=`, `>` and `>=` can be accelerated for `BitSigned`-`BigUnsigned` combinations of different widths. :robot: : Compare in the signed type when it can represent every value of the unsigned type, avoiding redundant sign checks. Share comparison logic across operand orders. Benchmarks on an Intel Core i7-1270P, pinned to a performance core, using Julia 1.14.0-DEV.3129 with LLVM 22.1.8. Both tables measure warmed 256-element arrays with Bool output. Scalar measurements disable LLVM loop and SLP vectorization. Timings exclude input generation but include loads, stores, and loop overhead. Values are changes in execution time, rounded to whole percentages. Negative means faster. `S` denotes the signed operand and `U` the unsigned operand. ### Scalar Loops | Operand types | S < U | U < S | S <= U | U <= S | |---|---:|---:|---:|---:| | Int16 / UInt8 | -44% | -44% | -42% | -45% | | Int32 / UInt8 | -45% | -44% | -44% | -44% | | Int32 / UInt16 | -45% | -43% | -43% | -43% | | Int64 / UInt8 | -43% | -43% | -41% | -43% | | Int64 / UInt16 | -43% | -43% | -42% | -44% | | Int64 / UInt32 | -42% | -42% | -41% | -41% | | Int128 / UInt8 | -32% | -34% | -34% | -25% | | Int128 / UInt16 | -27% | -32% | -35% | -28% | | Int128 / UInt32 | -26% | -33% | -35% | -25% | | Int128 / UInt64 | -27% | -35% | -38% | -25% | ### Vector Loops | Operand types | S < U | U < S | S <= U | U <= S | |---|---:|---:|---:|---:| | Int16 / UInt8 | -10% | -2% | -2% | 0% | | Int32 / UInt8 | 0% | +10% | +10% | +10% | | Int32 / UInt16 | 0% | +10% | +10% | +10% | | Int64 / UInt8 | -19% | -19% | -15% | -15% | | Int64 / UInt16 | -19% | -19% | -15% | -15% | | Int64 / UInt32 | -19% | -19% | -16% | -15% | | Int128 / UInt8 | -46% | -40% | -46% | -39% | | Int128 / UInt16 | -46% | -39% | -46% | -40% | | Int128 / UInt32 | -48% | -41% | -48% | -41% | | Int128 / UInt64 | -41% | -45% | -49% | -34% | Measurements compare averages of the fastest half of 3,001 samples, with 1,000 passes per sample and both benchmark orders checked. Identical-function controls were below 1% for all retained vector cases and 32 of 40 scalar cases. The remaining eight scalar controls differed by roughly 1–6%, so the percentages should be interpreted as approximate. The Int32 vector regressions result from LLVM choosing less efficient Boolean packing. Changing only that packing sequence removes the regression in an isolated assembly experiment. Reverting the operand-forwarding simplification produces identical vector assembly and therefore does not help. Overall, these workloads favor the optimization, although the benefit depends on operand types and execution mode. Ultimately, target-aware comparison optimization and Boolean packing belong in LLVM. * Make type conversion explicit Thanks for the suggestion, @vtjnash! --------- Co-authored-by: Patrick Häcker <patrick.haecker@bosch.com>
* REPL: Sync the ^C sweep test on the evaluation result In julia-pr build 2541 this test hung at the `Cancelled all in-flight work.` read and was hard-killed after the 900s timeout, on both i686-linux-gnu and aarch64-linux-gnu: https://buildkite.com/julialang/julia-pr/builds/2541#01a0aee0-fc2e-43c8-a079-4161fe774fa9 The sweep had in fact thrown - a `CancellationTokenSource` walk reached freed memory, since fixed by JuliaLang#63234 - and the keymap logs and swallows that, so the message never came and the read blocked forever. The test widened that window. Its stand-down step types `1` and reads until `julia> `, but the `^C` handler that armed the sweep ends in `transition(s, :reset); refresh_line(s)`, so a prompt is already sitting unread in the pipe and the read matches that instead, before `1` has been processed at all. The step meant to stand the arm down is never waited for, and everything after it races the evaluation. Looping the block under `-t1` hit the freed-memory error 5 times in 600 iterations as written, and once in 900 with the read below. Communicate through a result value, as the rest of the block does; the concatenation keeps the needle out of LineEdit's echo of the typed characters. Assisted-by: Claude Code (Opus 5) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * REPL: Reword the double-^C test comments Review of this test found its comments hard to read: they lean on invented vocabulary ("arms the sweep", "stands the arm down", "session epoch") instead of saying what each key press does. Describe the behaviour in plain terms, and spell out why the awaited results are written as concatenations. The marker expression is renamed to match. Assisted-by: Claude Code (Fable 5.1) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…uild build: fix USE_BINARYBUILDER=0 builds with sysroot toolchains
Member
|
Can you reopen the PR? Target |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream PR: mmtk/mmtk-core#1596