Skip to content

Update Ractor ivar isolation diagnostics for creator ownership - #1057

Closed
yaroslav-shopify wants to merge 1868 commits into
ractor-ivar-set-violation-messagefrom
ractor-ivar-set-violation-message-take-2
Closed

yaroslav-shopify wants to merge 1868 commits into
ractor-ivar-set-violation-messagefrom
ractor-ivar-set-violation-message-take-2

Conversation

@yaroslav-shopify

Copy link
Copy Markdown

Adapts the diagnostic improvement in ruby#17904 to creator-based class/module ownership, as suggested in Luke's review. Isolation errors identify the ivar and class/module while preserving the current ownership checks.

  • Rebase the original patch onto Ruby master at d41370aec5, keeping the original regression test and updating its expected message to “created by another Ractor”.
  • Include the ivar and class/module names in rb_ivar_get_at errors, matching rb_ivar_lookup.
  • Use rb_class_path when formatting names so a custom to_s is not invoked while constructing the isolation error.
  • Add coverage for owner access, rejected writes and deletion from the main Ractor to a worker-owned module, and consistent read diagnostics with an overridden to_s.

Validation: full make -j8 check passed with zero failures or errors, including 2,108 bootstrap tests, 37,006 Ruby tests, 33,344 Ruby specs, and the basic, tool, and SyntaxSuggest suites. Five focused ivar tests also passed with YJIT enabled.

Review note: the target branch still predates the rebase, so this branch-to-branch comparison includes upstream changes. The focused diff against the rebased master is limited to variable.c and bootstraptest/test_ractor.rb.

headius and others added 30 commits September 13, 2026 14:30
The functionality of this gem was moved into core Ruby starting
with Ruby 3.2, and at this point both the C extension and the Java
extension define no additional classes or methods. This patch
fully removes both extensions, replacing them with a single
io/wait.rb that warns about the gem deprecation. All other files
have been updated to reflect the removed extensions an deprecation.

Fixes ruby/io-wait#78

ruby/io-wait@4ae06a4278
Using the same technique from Daniel Lemire's blog post:
https://lemire.me/blog/2024/07/20/scan-html-even-faster-with-simd-instructions-c-and-c/)

This has multiple NEON code paths:

1. One for at least 64-bytes remaining in the input.
2. One for at least 16-bytes remaining in the input.

ruby/erb@f8314a7b86
A monitoring port received a bare :exited or :aborted, so a port could only
ever watch one ractor: with two, the token that arrived said what happened
but not to which.  Watching a group meant a port per ractor and a table
from the port back to the ractor -- which is what Ractor.select did, and
what anyone writing a supervisor had to do as well.

The token is now [ractor, :exited] or [ractor, :aborted], so one port can
watch a whole group and a supervisor needs nothing but monitor:

    workers.each { |r| r.monitor port }

    until workers.empty?
      r, status = port.receive
      workers.delete(r)
      workers << restart(r) if status == :aborted
    end

The pair is built by the receiver, from a basket type of its own, rather
than sent as a shareable array.  A shareable array would be simpler -- four
lines against thirty -- but it is left behind for every ractor that ever
terminates, and shareable objects are only reclaimed by a global
collection.  Creating 20,000 ractors and joining them:

                          wall       RSS held afterwards
    built on receipt     309 ms                    98 MB
    shareable array      570 ms                   533 MB

It cannot be sent as an ordinary copy either: the tokens go out from a
thread that has no VM stack left, which is what ractor_basket_new_ref
exists to avoid, and a copy pushes a tag.  Both halves are shareable
already, so the basket holds them as they are and the array is allocated
where it is read.

The array costs join, which received one symbol per call, 0.309us to
0.327us on a terminated ractor.  Ractor.select pays it back many times over
in the commit that follows.

This is an incompatible change to an experimental API, so NEWS says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
select watched each Ractor through a Port it made for that ractor, and kept
a Hash from the port back to the ractor because a bare exit token did not
say which ractor it was about.  Making a port keys it into the ractor's
port table and closing it takes it out again, and keying it into the Hash
assigned it an object_id, which for a T_DATA lives in a side table that
takes its synchronized path as soon as a second ractor exists -- always,
here.  All of that was paid per watched ractor.

Now that a token names its ractor, one port serves the whole call: select
monitors every watched ractor on it and reads the ractor out of the token.
When nothing but ports were passed there is still no port to make, and when
nothing but ractors were passed there is a single port to wait on, so the
selector is not built at all.

Per call, in microseconds of CPU, with none of the watched ractors ever
terminating so that what is measured is building and tearing down the wait
set rather than waiting:

    watched     before      after
          0       0.61       0.61
          1       1.53       1.16
          4       3.65       1.68
         16      12.38       3.66
         64      45.21      11.73
        128      91.10      22.58
        256     187.99      44.49

Per watched ractor, 0.732us to 0.171us.  Three repetitions alternating
which build ran first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
setup_gc_stat_symbols(), setup_gc_stat_heap_symbols() and the symbols
gc_info_decode() decodes were each filled on first use and guarded by
their own first element.  That element is also the first one written, so
a second ractor arriving while the first was still filling saw the guard
already set and skipped the setup, leaving the later entries at 0:

    Warning[:experimental] = false
    port = Ractor::Port.new
    256.times { Ractor.new(port) { |p|
      begin
        GC.stat(:heap_allocated_pages); p << :ok
      rescue ArgumentError => e
        p << e.message
      end
      nil
    } }
    p 256.times.map { port.receive }.reject { _1 == :ok }
    #=> ["unknown key: heap_allocated_pages"]

Reproduces about once per ten runs here; calling GC.stat once on the main
ractor first makes it go away, which is what pointed at the guard.

Fill all three tables from rb_gc_impl_init(), where no other ractor
exists yet, and drop the guards.  gc_info_decode()'s function-local
statics move to file scope so the new setup function can reach them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
heap_init_bytes is a slot count target applied per objspace, bytes divided
by each heap's slot size: heap_prepare fills a heap up to it before the
first collection, gc_sweep_finish_heap puts a floor under the free slots a
sweep should leave, and gc_marks_finish sums it across the heaps as the cap
on how much may stay free, which is what keeps empty pages around.  With
one shared objspace this 2.5 MB was paid once.  With per-Ractor GC every
Ractor's own objspace pays it again, and that is the steady heap a Ractor
holds: a Ractor that allocates 40,000 objects and discards them keeps
4.3 MB, where Ruby 4.0 kept 71 KB.

Add ractor_heap_init_bytes (RUBY_GC_RACTOR_HEAP_INIT_BYTES) and use it for
every objspace but the main one.  Its default, 0, is resolved at boot to
the smallest size that works: one slot in the largest heap.  Below that,
heap_prepare never forces a heap's first page and allocating there fails
with "cannot create a new page after GC", so the environment variable
rejects smaller values the way the other size parameters reject theirs.

Measured with 64 Ractors, each allocating 40,000 objects and then
discarding them, reported after a GC so that it is the heap floor rather
than live data:

    ruby                    kept per Ractor   GCs/Ractor    wall
    4.0.6                          70.9 KB        147.1   271 ms
    master                       4,310.3 KB          1.2   264 ms
    this commit                    302.2 KB         24.5   195 ms

Nothing between the minimum and one page's worth is worth choosing: from
1 KB to 16 KB the GC behaves identically (409 collections, 102 of them
major) and only the memory grows, 302 KB to 857 KB.  The 102 major
collections are all MAJOR_BY_NOFREE, and they stop at 64 KB, which is
HEAP_PAGE_SIZE: below one page's worth of slots, a sweep can never leave
enough free slots behind.  Removing them costs three times the memory
(910 KB) and buys 8% on a Ractor that allocates nothing but garbage, and
nothing at all once a Ractor holds a live set: with 200,000 objects live
the minimum runs 56 collections against 45 and takes 1% longer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A page with nothing live on it leaves its heap for the global empty pages,
and may be returned to the system from there.  When it is the heap's only
page the heap is left with none, and nothing grows a heap in that state:
gc_sweep_finish_heap's growth path asks for total_slots > 0, and
heap_prepare forces a first page only when it is not already sweeping.  An
allocation in that heap then reaches "cannot create a new page after GC".

It takes a heap whose page empties completely while another heap is being
swept, so a Ractor allocating a mix of sizes finds it and one allocating a
single size does not:

    Warning[:experimental] = false
    r = Ractor.new {
      20_000.times { |i| [Array.new(i % 300), "x" * (i % 2000), {a: 1}] }
    }
    r.value

Keep the page instead.  A heap that has been used once holds at least one,
which is what the rest of the code already assumes.  Recovering afterwards
in heap_prepare was measured too and holds the same 224-231 KB per Ractor,
but this keeps the invariant rather than repairing it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gc_marks_finish keeps at least heap_free_slots * r_mul slots free, where
r_mul was the number of Ractors allocating from this objspace, capped at
8.  Per-Ractor GC replaced that count with the VM-wide one, so every
Ractor's objspace now asks for free slots in proportion to how many
Ractors the whole VM has, although only its own Ractor allocates from it.

The floor can then exceed the heap it applies to.  A Ractor started with
the smallest initial heap holds about 7,200 slots; with one other Ractor
in the VM the floor is 4096 * 2, so no sweep can ever leave enough behind
and every collection escalates to a major one:

    1,000,000 allocations on one Ractor, all garbage

                        collections   major   reason
    before                      409     102   nofree
    after                       617       1   oldgen

Lowering RUBY_GC_HEAP_FREE_SLOTS to 1024 has the same effect, which is
what identified the floor.  The major collections free nothing: they cost
8% on a Ractor that allocates nothing but garbage and nothing once one
holds a live set, and the heap does not shrink.  Memory is unchanged
either way, at 298 KB per Ractor.

The main objspace keeps the VM-wide count it has used since before
per-Ractor GC.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The initial heap size is a slot count target, bytes / slot_size for each
heap, so below one slot in the largest heap heap_prepare never forces that
heap's first page.  RUBY_GC_RACTOR_HEAP_INIT_BYTES rejects such values;
RUBY_GC_HEAP_INIT_BYTES accepts anything above zero, and dies partway
through boot:

    $ RUBY_GC_HEAP_INIT_BYTES=512 ./ruby -e 'puts "ok"'
    encdb.so: [BUG] cannot create a new page after major GC

Give it the same bound.  Too small a value is now ignored with a warning
under -w, as one that overflows or does not parse already is:

    RUBY_GC_HEAP_INIT_BYTES=512 (default value: 2621440) is ignored
    because it must be greater than 1023.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It's failing compilation on `rmv7a-linux-androideabi30-clang`:

```
escape.c:223:34: error: call to undeclared function 'vqtbl1q_u8'; ISO C99 and later do not support implicit function declarations [-Wimplicit-function-declaration]
  223 |     const uint8x16_t looked_up = vqtbl1q_u8(escape_char_by_low_nibble, low_nibbles);
      |                                  ^
escape.c:223:34: note: did you mean 'vtbl1_u8'?
/home/chkbuild/opt/android-ndk-r29/toolchains/llvm/prebuilt/linux-x86_64/lib/clang/21/include/arm_neon.h:33478:48: note: 'vtbl1_u8' declared here
 33478 | __ai __attribute__((target("neon"))) uint8x8_t vtbl1_u8(uint8x8_t __p0, uint8x8_t __p1) {
       |                                                ^
escape.c:223:22: error: initializing 'const uint8x16_t' (vector of 16 'uint8_t' values) with an expression of incompatible type 'int'
  223 |     const uint8x16_t looked_up = vqtbl1q_u8(escape_char_by_low_nibble, low_nibbles);
      |                      ^           ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
escape.c:236:14: error: call to undeclared function 'vpaddq_u8'; ISO C99 and later do not support implicit function declarations [-Wimplicit-function-declaration]
  236 |     folded = vpaddq_u8(folded, folded);
      |              ^
escape.c:236:12: error: assigning to 'uint8x16_t' (vector of 16 'uint8_t' values) from incompatible type 'int'
  236 |     folded = vpaddq_u8(folded, folded);
      |            ^ ~~~~~~~~~~~~~~~~~~~~~~~~~
escape.c:237:12: error: assigning to 'uint8x16_t' (vector of 16 'uint8_t' values) from incompatible type 'int'
  237 |     folded = vpaddq_u8(folded, folded);
      |            ^ ~~~~~~~~~~~~~~~~~~~~~~~~~
escape.c:238:12: error: assigning to 'uint8x16_t' (vector of 16 'uint8_t' values) from incompatible type 'int'
  238 |     folded = vpaddq_u8(folded, folded);
      |            ^ ~~~~~~~~~~~~~~~~~~~~~~~~~
escape.c:256:25: error: call to undeclared function 'vpaddq_u8'; ISO C99 and later do not support implicit function declarations [-Wimplicit-function-declaration]
  256 |     uint8x16_t folded = vpaddq_u8(vpaddq_u8(t0, t1), vpaddq_u8(t2, t3));
      |                         ^
escape.c:256:16: error: initializing 'uint8x16_t' (vector of 16 'uint8_t' values) with an expression of incompatible type 'int'
  256 |     uint8x16_t folded = vpaddq_u8(vpaddq_u8(t0, t1), vpaddq_u8(t2, t3));
      |                ^        ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
escape.c:257:12: error: assigning to 'uint8x16_t' (vector of 16 'uint8_t' values) from incompatible type 'int'
  257 |     folded = vpaddq_u8(folded, folded);
      |            ^ ~~~~~~~~~~~~~~~~~~~~~~~~~
```

ruby/erb@45b481473c
(ruby/stringio#225)

`Arrays.fill(byte[], int, int, byte)` takes **fromIndex, toIndex**, but
`ungetbyteCommon` passes `memset`'s **destination, length**:

```java
if (rest > cl) Arrays.fill(strBytes, len, rest - cl, (byte) 0);
```

straight from `ext/stringio/stringio.c:1181`, `memset(s + len, 0, rest -
cl)`. It also drops the ByteList `begin` offset that the C `s +`
supplies.

Pushing back onto a StringIO positioned more than the pushback length
past the end of its string then throws instead of zero-filling:

```ruby
s = StringIO.new("abc"); s.pos = 5; s.ungetbyte("d")
# JRuby:  Java::JavaLang::IllegalArgumentException: fromIndex(3) > toIndex(1)
# CRuby:  s.string == "abc\0d", s.pos == 4
```

The two siblings in this file already spell the correct form —
`truncate` (`StringIO.java:1462`) uses `buf.getBegin() + plen,
buf.getBegin() + l`, and `strioExtend` (`:1044-1048`) uses `begin +
olen, begin + pos`. This one line is the odd one out, so the fix reuses
their shape rather than adding anything.

The added expectations extend the existing `test_ungetc_padding` /
`test_ungetbyte_padding`, which stop one short: both start from
`StringIO.new()`, so `len == 0` makes the wrong call accidentally
correct.

<details>
<summary>Verification</summary>

Both backends built from this clone, `$LOADED_FEATURES` printed and
checked on every row.

A differential sweep of 2100 cases (7 base strings x 25 positions x 6
pushback values x `ungetbyte`/`ungetc`), JRuby against the C extension:

| build | rows differing from CRuby |
|---|---|
| current `master` | **132** |
| with the fix | **0** |

Every one of the 132 is the `IllegalArgumentException`.

| row | result |
|---|---|
| the two padding tests on `master` | 2 errors |
| with the fix | 2 tests, 12 assertions, 0 failures, 0 errors |
| revert | 2 errors |
| widen the guard to `rest > 0` | 1 error |
| `fromIndex` one low | 1 failure, 1 error |
| `toIndex` one high | passes — see below |

That last mutant survives, and I checked whether it was a gap in my test
or genuinely equivalent: the extra byte lands at `s + pos` after `pos -=
cl`, which `System.arraycopy(..., strBytes, s + pos, cl)` two lines
later overwrites unconditionally. Sweeping it over the same 2100 cases
gives **0 differences from the fixed build**, so it is an equivalent
mutant rather than an uncovered line, and I did not add a row for it.

Full `test/stringio/test_stringio.rb`: CRuby 3.2 **104 tests, 646
assertions, 0 failures, 0 errors**; JRuby 10.0.6 **104 tests, 643
assertions, 0 failures, 0 errors**.

Every added expectation was run against the C extension first, so it
encodes CRuby behaviour rather than my reading of it.

Two things I want on the record. The `s +` half is not separately
covered — I could not construct a StringIO whose backing ByteList has
`begin != 0` through the public API. It is included because both
siblings and the C both require it; if you would rather keep the change
minimal, the two-index correction alone fixes the crash. And
`ruby/.github`'s SECURITY.md routes vulnerabilities privately: this is
the JVM's own bounds check firing rather than memory unsafety, and I
could not produce any stale-byte disclosure, so I filed publicly — say
the word if you would rather it had gone the other route.

The CI matrix runs `jruby-head`; the closest I could obtain is JRuby
10.0.6. Note `jruby:9.4` cannot compile `ext/java` at master at all
(`Helpers.memchr` signature mismatch), so 10.x is already the floor.
</details>

---

Disclosure: I used Claude (an AI assistant) while preparing this change.
Every result above I ran and verified myself.

ruby/stringio@9d3e502b3a
This commit changes Enuemrator::Lazy#uniq to use a temporary Set instead
of Hash for performance and memory savings. We can see in the following
benchmark:

    a = ((0...10_000).to_a + (0...10_000).to_a).lazy
    1_000.times { a.uniq.to_a }

Results:

    Benchmark 1: master
      Time (mean ± σ):     867.5 ms ±  13.6 ms    [User: 857.1 ms, System: 7.8 ms]
      Range (min … max):   853.9 ms … 899.4 ms    10 runs

    Benchmark 2: branch
      Time (mean ± σ):     801.4 ms ±   7.1 ms    [User: 791.3 ms, System: 7.4 ms]
      Range (min … max):   789.4 ms … 810.7 ms    10 runs

    Summary
      branch ran
        1.08 ± 0.02 times faster than master
Bumps the github-actions group with 2 updates in the / directory: [actions-rust-lang/setup-rust-toolchain](https://github.com/actions-rust-lang/setup-rust-toolchain) and [taiki-e/install-action](https://github.com/taiki-e/install-action).


Updates `actions-rust-lang/setup-rust-toolchain` from 1.17.0 to 2.0.0
- [Release notes](https://github.com/actions-rust-lang/setup-rust-toolchain/releases)
- [Changelog](https://github.com/actions-rust-lang/setup-rust-toolchain/blob/main/CHANGELOG.md)
- [Commits](actions-rust-lang/setup-rust-toolchain@166cdcf...ecabd13)

Updates `taiki-e/install-action` from 2.87.7 to 2.87.8
- [Release notes](https://github.com/taiki-e/install-action/releases)
- [Changelog](https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md)
- [Commits](taiki-e/install-action@84f5ac3...d438492)

---
updated-dependencies:
- dependency-name: actions-rust-lang/setup-rust-toolchain
  dependency-version: 2.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: github-actions
- dependency-name: taiki-e/install-action
  dependency-version: 2.87.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps the github-actions group with 5 updates in the / directory:

| Package | From | To |
| --- | --- | --- |
| [zizmorcore/zizmor-action](https://github.com/zizmorcore/zizmor-action) | `0.6.3` | `0.6.4` |
| [github/codeql-action/init](https://github.com/github/codeql-action) | `4.37.9` | `4.38.0` |
| [github/codeql-action/analyze](https://github.com/github/codeql-action) | `4.37.9` | `4.38.0` |
| [github/codeql-action/upload-sarif](https://github.com/github/codeql-action) | `4.37.9` | `4.38.0` |
| [taiki-e/install-action](https://github.com/taiki-e/install-action) | `2.87.8` | `2.87.11` |



Updates `zizmorcore/zizmor-action` from 0.6.3 to 0.6.4
- [Release notes](https://github.com/zizmorcore/zizmor-action/releases)
- [Commits](zizmorcore/zizmor-action@70fb788...cc914d7)

Updates `github/codeql-action/init` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@cdf488f...b96794f)

Updates `github/codeql-action/analyze` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@cdf488f...b96794f)

Updates `github/codeql-action/upload-sarif` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@cdf488f...b96794f)

Updates `taiki-e/install-action` from 2.87.8 to 2.87.11
- [Release notes](https://github.com/taiki-e/install-action/releases)
- [Changelog](https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md)
- [Commits](taiki-e/install-action@d438492...9534c84)

---
updated-dependencies:
- dependency-name: zizmorcore/zizmor-action
  dependency-version: 0.6.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
- dependency-name: github/codeql-action/init
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-actions
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-actions
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-actions
- dependency-name: taiki-e/install-action
  dependency-version: 2.87.11
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
...

Signed-off-by: dependabot[bot] <support@github.com>
308677c moved the exit tokens over to ractor_basket_new_exit(), which
left this one with no caller.  Its comment carried the reason those tokens
cannot push a tag, so that moves to ractor_basket_new_exit().

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Builtin methods written in Ruby are defined in the master box, so the
box resolution stops at their frames and cannot see classes defined in
the caller's box. The new attribute makes such methods operate on the
caller's box, as CFUNC frames already do.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Marshal.load is a builtin Ruby method, so with Ruby::Box enabled it
looked up dumped class names in the master box and failed even for a
same-process round-trip. Mark it :caller_box to resolve the class
names in the caller's box, which Marshal.dump (a C function) already
does.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The box of the immediately previous frame can be the root box, e.g. when
a proc made in the root box calls Marshal.load, and the classes of the
caller's data are invisible there. Walk the caller frames up to the
nearest user box instead, and rename the attribute after what it finds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Using the same technique from Daniel Lemire's blog post:
https://lemire.me/blog/2024/07/20/scan-html-even-faster-with-simd-instructions-c-and-c/)

This has multiple NEON code paths:

1. One for at least 64-bytes remaining in the input.
2. One for at least 16-bytes remaining in the input.

ruby/erb@9717249de6
* `tr_trans_pair` is now responsible for consuming the matches,
akin to the ERB change in 4ca48a1.
This makes it impossible to shift by 64 or more, hence removes a
branch in `next_match`. Also ensure `s` and `matches_bitmap` members
stays in sync.

* Add missing `needles_count` check in `tr_trans_pairs_search_sse2`.

Reported-By: Ashley Allen <a.allen3@herts.ac.uk>
Reported-By: Asten Rooky
BurdetteLamar and others added 28 commits September 21, 2026 11:14
String#tr can crash if keys are removed from the translation hash since
we stack allocate a buffer pairs. If entries are deleted during runtime,
then we won't fill the pairs buffer which can crash because it will be
reading uninitialized values out of pairs.

The following script demonstrates the crash:

    h = {}
    obj = Object.new
    obj.define_singleton_method(:to_str) do
      h.clear
      "a"
    end
    h[obj] = "x"
    ("b".."z").each { |c| h[c] = (c.ord + 1).chr }
    puts "ab".tr!(h)
fu_copy_stream0 has had no caller since b26846bf replaced both uses with direct IO.copy_stream calls in 2008. The retained private wrapper is unreachable.

ruby/fileutils@feba38f651
…able

In user-namespace environments (e.g. ChromeOS Crostini) getgroups(2) can report supplementary groups such as the overflow GID 65534 (nobody) that a non-root process cannot actually chgrp a file to. TestFileUtils#setup built @groups straight from `[Process.gid] | Process.groups`, so the chown tests tried to chgrp to such a group and failed with EPERM instead of being skipped.

Filter @groups with a one-time capability probe down to the groups the process can actually assign, so the affected tests skip via their existing `return unless @groups[1]` guard when there is no usable second group.

ruby/fileutils@8dc0266054

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Skip the test_rm_r_no_permissions test under the root user, as deletion always succeeds.

Signed-off-by: Jiaying Song <jiaying.song.cn@windriver.com>

ruby/fileutils@3c831389c5
Since `Process.uid` returns `0` always on mingw or mswin, `root_in_posix?` returns `true` too.

ruby/fileutils@d18fbf0c56
If the source string is modified when converting offset to an integer,
then there could be an out-of-bounds access because the length of the
string is captured before to_int is called. For example, the following
script triggers an ASAN error:

    s = "héllo" * 1_000_000
    obj = Object.new
    obj.define_singleton_method(:to_int) do
      s.replace("é" * 1000)
      450000
    end
    p s.byteindex("l", obj)
When the array is shrunk during rb_arithmetic_sequence_beg_len_step,
ary_subseq_len returns -1. Return an empty array instead of passing the
negative length to ary_make_partial or ary_make_partial_step.

[Bug #22325]
We need to keep track of pending pages in RubyHeapTrigger otherwise a large
allocation may cause infinite number of GCs to be ran since it will always
appear that we have enough space in the heap (and thus the heap won't grow)
but still not have enough for the allocation.

ruby/mmtk@dd8034050d
It's safe to free these types during Ractor-local GC in Ruby 4.1. We've
seen the freeing of this type show up in profiles before.

ruby/openssl@e481c6408f

Co-authored-by: Kazuki Yamaguchi <k@rhe.jp>
Ref: https://bugs.ruby-lang.org/issues/17593.

Because `__FILE__` is currently considered a static literal,
it's entirely evaluated during compilation.

The problem is that ISeqs can be serialized and reloaded later,
but since `__FILE__` is entirely computed during compilation,
a serialized ISeq is coupled with the location of the source code.

This prevent compiling the the iseq with the source in one path,
and loading it later with the source in another path, which prevent
bootsnap precompilation from working with some builds systems.

I'm working on changing the compiler so that `__FILE__` is no longer
a static value, but a hidden methodthat returns the current ISeq path.

ruby/prism@f16d852c8c
The conversion of the pattern to string can cause the string to be modified.
This can cause a buffer overflow since it uses the original length of the
string. For example, the following script reports a buffer overflow with
ASAN:

    s = "héllo" * 1000
    obj = Object.new
    obj.define_singleton_method(:to_str) do
      s.replace("hé")
      "l"
    end
    p s.rpartition(obj)
Bumps the github-actions group with 5 updates in the / directory:

| Package | From | To |
| --- | --- | --- |
| [ruby/setup-ruby](https://github.com/ruby/setup-ruby) | `1.323.0` | `1.324.0` |
| [github/codeql-action/init](https://github.com/github/codeql-action) | `4.38.0` | `4.38.1` |
| [github/codeql-action/analyze](https://github.com/github/codeql-action) | `4.38.0` | `4.38.1` |
| [github/codeql-action/upload-sarif](https://github.com/github/codeql-action) | `4.38.0` | `4.38.1` |
| [taiki-e/install-action](https://github.com/taiki-e/install-action) | `2.87.13` | `2.87.16` |



Updates `ruby/setup-ruby` from 1.323.0 to 1.324.0
- [Release notes](https://github.com/ruby/setup-ruby/releases)
- [Changelog](https://github.com/ruby/setup-ruby/blob/master/release.rb)
- [Commits](ruby/setup-ruby@984c0c8...a0102e0)

Updates `github/codeql-action/init` from 4.38.0 to 4.38.1
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@b96794f...1c5b675)

Updates `github/codeql-action/analyze` from 4.38.0 to 4.38.1
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@b96794f...1c5b675)

Updates `github/codeql-action/upload-sarif` from 4.38.0 to 4.38.1
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@b96794f...1c5b675)

Updates `taiki-e/install-action` from 2.87.13 to 2.87.16
- [Release notes](https://github.com/taiki-e/install-action/releases)
- [Changelog](https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md)
- [Commits](taiki-e/install-action@26e9283...9114bf4)

---
updated-dependencies:
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.38.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
- dependency-name: github/codeql-action/init
  dependency-version: 4.38.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.38.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
- dependency-name: ruby/setup-ruby
  dependency-version: 1.324.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-actions
- dependency-name: taiki-e/install-action
  dependency-version: 2.87.15
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Calling to_int on the count may run Ruby code that modifies the array.
This can cause an out-of-bounds in Array#permutation. For example, the
following code crashes:

    ary = (1..1_000).to_a
    obj = Object.new
    obj.define_singleton_method(:to_int) do
      ary.replace(Array.new(5) { :x })
      2
    end
    ary.permutation(obj) { }
Show that dangling symbolic links return false and clean up the
example files.

Follow up on ruby#18571.
Also simplify the notation of the special const reserved words.
If the Dir object is closed in the block of Dir#each/each_child/scan, then
it will crash because dirp->dir will be a NULL pointer. The following
script demonstrates the crash:

    d = Dir.open("/")
    d.each { d.close }
- ### Problem

  When a Ractor isolation error happens due to setting an ivar on a
  class, the message doesn't include the name of the ivar and neither
  the name of the class. It makes it difficult to track down where the
  isolation problem comes from.

  I'd like the message to be similar as what we get when a Ractor
  access an unshareable ivar.

  This is the current behaviour

  ```ruby
  class Foo
    def self.foo
      @foo ||= "test"
    end
  end

  Ractor.new do
    Foo.foo # can not set instance variables of classes/modules by non-main Ractors
  end.join

  Foo.foo

  Ractor.new do
    Foo.foo # can not get unshareable values from instance variables of classes/modules from non-main Ractors (@foo from Foo)
  end.join
  ```

  ### Solution

  Modify the error message to include the ivar name and the class/module name.
Include ivar and class names in JIT read errors, avoid invoking custom to_s while formatting isolation errors, and cover creator ownership across Ractors.

Validation: make -j8 check (passed).
@yaroslav-shopify
yaroslav-shopify deleted the ractor-ivar-set-violation-message-take-2 branch September 22, 2026 07:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.