Conversation
ygao-g
commented
Aug 27, 2026
Comment on lines
+265
to
+271
| # Plain DNS64 synthesizes only for names with no AAAA, and the names that | ||
| # matter here have real AAAA records pointing at addresses a host with no | ||
| # IPv6 egress cannot reach. Only translate_all forces them through the | ||
| # prefix -- which is why this is opt-in: where IPv6 egress does work it | ||
| # replaces reachable answers with unreachable ones. | ||
| # | ||
| # translate_all cannot share a server block with the cluster zones. dns64 |
Owner
Author
There was a problem hiding this comment.
Move these comments to before the corresponding logic.
Owner
Author
There was a problem hiding this comment.
this applies from 265 to 283
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
August 27, 2026 20:28
a22c540 to
f3e7447
Compare
Adds ate-setup, a Go-implemented version of the hack scripts
This was done with heavily machine-translated code iterated
with testing.
…ds (agent-substrate#1276) The Capabilities, SecurityContext, ImageVolumeSource and Volume.image fields landed without doc comments, so apitool's documented rule fails on them and they are not in the exemption backlog. Document them rather than growing the backlog. Fixes failure on main > It's a good idea to open an issue first for discussion. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
agent-substrate#232 Add documentation guide for CSI volumes in Substrate - [X ] Tests pass - [ X] Appropriate changes to documentation are included in the PR
…gent-substrate#1263) The counter tutorial's headline claim — state preserved across suspend/resume — is false on the default path: with `onCommit: Data`, suspend excludes process memory, so the in-memory counter resets while only the durable file counter continues. **Reproduced live** (GKE, template config identical to main): memory 5 → **1**,2,3 after suspend/resume; suspend manifest `scope:"data"`, `pages.img` 71 B (vs 2.2 MB golden Full). The micro-VM variant has the same broken promise via a different path: `fromData: Golden` restarts the count from the golden state — the e2e Golden case itself asserts memory=1 after suspend. **Fix** (per maintainer decision — config *and* docs, both variants matching): `onCommit: Full` on both counter templates, drop the now-dead `onResume` block on the micro-VM one, and make the README precise about which scope preserves what. **Verified live**: with `Full`, memory count runs 5 → **6**,7,8 across suspend/resume; snapshot manifest `scope:"full"`, `pages.img` 2.4 MB. No coverage lost: `TestDurableDirLifecycle` exercises every `onCommit`/`onPause`/`fromData` combination (including Data and Golden) with its own parameterized templates, independent of the demo defaults — and its `Full/Full` case already asserts the memory count continues, in both sandbox classes. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
August 27, 2026 23:31
f3e7447 to
c711eca
Compare
…bstrate#1264) Create/Suspend/Resume Actor workflows now uses the `Actor.actor_template` ObjectRef when specified, otherwise falls back to the CRD in cmd/ateapi/internal/controlapi/template_convert.go.
atelet's spec was written for runsc: a hardcoded `runsc` hostname, CRI pause/sandbox annotations and `dev.gvisor.spec.mount.*` hints. The micro-VM shim compensated by rewriting every spec at runtime (ensureKataCompatibleSpec), swapping the whole mount set and rebuilding it from one hand-written builder per volume kind — so a volume kind atelet learned about reached the guest only once the shim was taught about it too. atelet now emits a spec that names no runtime, and each ateom shapes it for the runtime it drives. Whatever atelet adds therefore reaches both, and a bind the micro-VM shaper cannot place in the guest is an error rather than a silent drop. Fixes agent-substrate#709. > It's a good idea to open an issue first for discussion. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
…rkloads (--mem-targetLarge mem bench (agent-substrate#1130) ## What this PR does: Adds large-memory benchmark suites: glutton actors that hold a resident 1–2Gi working set, so the suspend/resume path is measured at realistic application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm}, tracked. The boomer GluttonUser fills each actor to a configured target via chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size field is int32) after resume and **before the first suspend**, so every snapshot from cycle one onward carries the full working set. `WriteRAM` writes incompressible random bytes, so snapshots are genuinely target-sized rather than zstd-compressing away (a zero-filled working set compresses ~35:1 and would shrink the upload/download phases to a few MB). ## How it's configured The target flows through the established boomer dynconfig channel, like the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` → `/boomer-config` → the Go worker. Because it's per-user-class runtime config, heterogeneous and changing workload shapes are expressible with no redeploy, and the deploy stack needs no changes. Changes: - `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it was gRPC-only, unreachable in `--mode=http` deployments) - `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user cycle — runs once per actor (glutton holds the allocations across suspend/resume), retried on failure, reported as its own `GluttonFillRAM` stats row so it never pollutes ping/resume numbers - `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`: new `mem_target_bytes` knob - `tests.yaml`: the four tracked suites, with `actorMemory` sized above the target for headroom ## Behavior notes - The golden template snapshot and each actor's cold boot stay small: the working set exists from the first fill onward. All steady-state suspend/resume cycles measure at size; `ResumeActorColdStart` does not. - Each actor's first suspend uploads the first at-size snapshot and follows the fill — cycle one is an expected outlier on the dashboards. - A mid-run change to `mem_target_bytes` applies to newly spawned users; already-filled actors keep their size, so each actor's cycles stay comparable. - Existing suites are untouched: with no flag the target is 0 and the fill is a no-op. ## Testing - `go test -race` across `cmd/benchmarking/glutton`, `internal/benchmarking/boomer/...`, and the glutton fake: PASS. New tests cover chunking to an exact target, disabled-by-default, and fail-then-retry. - Cluster verification (microvm, 1Gi target): verified that the RAM fill completes before the initial suspend, snapshot upload reflects the expected ~1Gi payload, and steady-state suspend/resume cycles succeed without error.
…t-substrate#1279) agent-substrate@4809097#diff-51a595730fdeea0b1efc0394d1b3e866426a9acb000b6f2110262867b9aaeae7 added exemptions unrelated to the PR agent-substrate#1276 removed the need for exemptions, which were currently failing on main This PR cleans up the now failing defunct exemptions.
That's twice today that we missed regenerating the proto clients: 1. mid-PR on agent-substrate#1276 (caught by code review) 2. agent-substrate#1130 (missed and still missing) So this PR does: 1. regenerate with latest after agent-substrate#1130 2. ensure verify will catch this (so CI should fail if we haven't run it), and that update will run it when developing fixes agent-substrate#738 - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
On a fresh GCP project nothing documented which IAM bindings atelet needs: setup-gcp creates them silently, and anyone who cannot run the tool with project-level IAM permissions - or needs to audit what it did - had to read cmd/iam.go and cmd/bucket.go. Spell out the exact members, roles, and resources, the Workload Identity prerequisites, and the gcloud equivalents, and point the README quickstart at it.
certificates.k8s.io/v1beta1 podcertificaterequests and clustertrustbundles cannot be enabled in place on an existing cluster - the update is accepted but the APIs never become served, and the install hangs waiting for ClusterTrustBundles. Warn in the create cluster docs, show the bring-your-own-cluster flag, and name the symptom.
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
August 28, 2026 16:23
c711eca to
1cb5a78
Compare
…#1286) When checking documentation, remove `validation-gen` tags so we don't have false negative results if a field *only* has tags and no documentation.
…ate resource (agent-substrate#1281) * Added kubectl ate client changes for applying the counter demo. * Added demo in both gVisor and microVM. * Updated documentation. Example output: ``` $ kubectl ate get actortemplates -a ate-demo-counter-substrate ATESPACE NAME SANDBOX CLASS STATUS AGE ate-demo-counter-substrate counter SANDBOX_CLASS_GVISOR Ready 43s $ kubectl ate get actortemplate counter -a ate-demo-counter-substrate -o yaml actorTemplates: - containers: - command: - /ko-app/counter - --extra-port=9090 - --tcp-port=9091 image: gcr.io/zoezhao-gke-dev/ate-images/counter-7b2c368808ac33f45c7ab87955715526@sha256:9eb06b137bfd9f74f7ffa7c6052bb189a7b448801da81931a2124600a2622fa5 name: counter readyz: httpGet: path: /readyz port: 80 volumeMounts: - mountPath: /home/counter name: data metadata: atespace: ate-demo-counter-substrate createTime: "2026-08-28T02:09:45.981177255Z" name: counter uid: a8520ce2-f1f5-45d9-8d94-a436a72c5bf1 updateTime: "2026-08-28T02:09:56.512455175Z" version: "3" resources: limits: - name: cpu quantity: "1" - name: memory quantity: 512Mi sandboxConfig: configName: gvisor-default sandboxClass: SANDBOX_CLASS_GVISOR snapshotsConfig: onCommit: SNAPSHOT_CONTENT_SCOPE_FULL onPause: SNAPSHOT_CONTENT_SCOPE_FULL storageLocation: gs://snapshot-substrate-test-zoezhao-gke-dev/ate-demo-counter-substrate/ status: goldenSnapshotStatus: goldenSnapshot: atespace: ate-demo-counter-substrate name: 208d6f11-92e7-4ac0-9cab-74cd0b8a7992 takeGoldenSnapshotAt: "2026-08-28T02:09:56.041417048Z" volumes: - durableDir: {} name: data type: DurableDir workerSelector: matchLabels: workload: counter-substrate ``` More testing log in gpaste/5520785792434176
Flesh out the agentgateway implementation as a router and egress PEP. This also changes the atunnel CONNECT port to do 8443. Most of this is agentgateway contained, but the one other system we touch is localca; agentgateway's MITM support requires the TLS cert and key in the secret instead of the JSON pool (we don't rely on sdsmint). - [X] Tests pass - [X] Appropriate changes to documentation are included in the PR --------- Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
…rs, behind a flag
…n) (agent-substrate#1280) #### What this pr do Re-dirties part of the glutton working set on every benchmark cycle, so repeated suspends snapshot an actor whose memory is changing — like a live application's — instead of a set that was filled once and never touched again. Each iteration (after the one-time fill), the GluttonUser sends one `WriteRAM` request with `WRITE_MODE_OVERWRITE`, re-randomizing the first `mem_churn` bytes of the working set in place. Overwrite mode already existed on the glutton API; this PR adds no proto changes and no glutton server changes — it is driver + config wiring only. This is the follow-up from agent-substrate#1130's: the original self-driving memload continuously re-dirtied its pages, and that property was dropped in the move to API-driven fill. This restores it in the API-driven shape with a dialable amount instead of the old all-or-nothing full pass. #### Why it matters A fill-once working set is static: every suspend after the first packs up identical memory. If snapshotting ever gains incremental / dirty-page optimizations, a static benchmark would measure almost nothing from cycle two onward — and would score "upload nothing" as an infinite win. With churn, every cycle carries a known amount of freshly-dirtied pages, and the knob is sweepable (64Mi, 256Mi, …) so a future differential-snapshot optimization can be demonstrated as "upload cost scales with churn size, not total size."
…rate#1287) Contains the gVisor fixes for agent-substrate#1090 agent-substrate#992
…strate#1210) Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the constraints so the field can be left unset. This is the basis for let an empty workerImage lets the controller inject a versioned default image chosen by the pool's sandbox class. Part of agent-substrate#861 - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
…strate#1285) The networking suite's direct-access and arbitrary-port tests build their actors from the counter fixture and behave the same whichever egress gateway is deployed, so re-running them after the sdsmint swap adds CI time without adding coverage. Only egressFixture() reads E2E_EGRESS_MITM, so restrict both MITM lanes to the tests that use it. > It's a good idea to open an issue first for discussion. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Necessary for agent-substrate#1266, but functionally useful regardless. This eliminates the MPL go-reap dependency in favor of our own implementation tuned to our needs. Our package has utilities for spawning child processes without holding a RWMutex to avoid convoying lots of concurrent activations in the future.
Add maxItems checks to lists and simplify uniqueness checks.
…bstrate#595) DNS lookups are case-insensitive (RFC 4343), so a client that resolves `MyActor.MySpace.actors.resources.substrate.ate.dev` reaches the router with that spelling preserved in `Host`/`:authority` (Go's http.Client and `curl -H` both send the host as written). `ParseActorDNSName` compared it byte for byte against a strictly lower-case suffix and name pattern, so the request was rejected with a 404 by `invalidHostErr` even though its DNS lookup had resolved to that very actor. Fold the name to lower case before parsing. Actor and atespace names are always lower case (`ResourceNameRegexPattern`), so a folded name either addresses the same actor its lookup resolved to, or fails validation exactly as before.
…#1300) Remove `ActorTemplateStatus.sandbox_assets` along with the `SandboxAssets` messages it referenced. Nothing ever wrote the field: sandbox assets are resolved from the WorkerPool and SandboxConfig objects at resume time and travel to atelet via ateletpb, so freezing them into the template status never materialized. This requires more thought, one option is to add it as a field of GoldenSnapshotStatus in the future, and make GoldenSnapshotStatus a repeated field to allow multiple goldens.
…#416) Part of agent-substrate#207, following the plan discussed there with @juli4n and @kannon92: plumbing plus the non invasive fixes first, with the existing invasive findings excluded so new APIs do not regress (as @kannon92 suggested in agent-substrate#207 (comment)). The linter itself was originally suggested by @BenTheElder in agent-substrate#188. Plumbing: - hack/tools/kube-api-linter module with the golangci-lint-kube-api-linter tool, same pattern as the other tools. - .golangci-kal.yaml config. The pre-existing findings from rules that need Go API changes (nomaps, nonpointerstructs, nophase, optionalfields, requiredfields, ssatags) are excluded rather than disabled, and each exclusion rule names the fields it excuses, so the rules still apply to new API files and to new fields on the existing types. Fixing them on the existing types is the follow up in agent-substrate#207. - hack/verify/kube-api-linter.sh runs it against ./pkg/api/... and is picked up by hack/verify-all.sh in CI. Fixes in pkg/api/v1alpha1 (doc and marker only, no Go API change): - commentstart: field godocs now start with the serialized field name. - conditions: Conditions moved to the first position in ActorTemplateStatus and got the listType, listMapKey and patch markers. - defaultorrequired: SandboxConfigSpec.SandboxClass had both a default and required; now optional with omitempty, matching the same field on WorkerPoolSpec and ActorTemplateSpec. The ValidatingAdmissionPolicy in sandboxconfig-validation.yaml is unaffected since CRD defaulting runs before admission, so spec.sandboxClass is always set by then. - optionalorrequired: missing +optional markers added on ActorTemplateStatus fields. - defaults: configured preferredDefaultMarker to kubebuilder:default since CRDs are generated with controller-gen. Generated CRDs regenerated. Structural schema changes are only the conditions listType/listMapKey and sandboxClass no longer in the required list (it is defaulted, so behavior is the same); the rest is description text. Note: my earlier count of 50 issues in agent-substrate#207 was capped by the default golangci-lint issue limit. With the cap removed the real total is 86; the extra ones are requiredfields (20) and ssatags (4), both in the excluded set above. Test plan - hack/verify/kube-api-linter.sh passes (0 issues). - Exclusions verified both ways: dropping them brings the pre-existing findings back, and a scratch field added to WorkerPoolSpec is still reported (optionalfields), which the earlier per-file exclusion swallowed. - TestSandboxConfigValidation passes against the regenerated CRD. - go build ./... and go test ./pkg/api/... ./cmd/atecontroller/... pass. - gofmt, shellcheck and the regular golangci-lint on pkg/api are clean.
ateapi is a headless Service, so its DNS name resolves to one address per replica and the client picks. This one asked for no policy, which means gRPC's pick_first: it connects to whichever address answers first and sends everything there for the life of the connection. So a client talks to one replica however many are running, and adding replicas moves no load. Measured with eight load generators against two replicas, one took 97% of the traffic. internal/ateapiauth, which atelet and the controllers dial through, has asked for round_robin all along. This is the same policy for the clients that do not. Pulled out from agent-substrate#1266, I ran into this while doing some custom benchmarking using ateclient, but this is already a bug for kubectl-ate, demos, ... - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR AI-assisted.
@BenTheElder @zlammerts-svg Thanks again for driving this Tim!
…-substrate#1060) Part of agent-substrate#246 The egress Envoy pinned `dns_lookup_family` to `V4_ONLY` in both the dynamic forward proxy filter and the dynamic forward proxy cluster. On an IPv6-only cluster every upstream connection failed: Envoy reported a DNS resolution failure and the actor got a 503. Seven sites across the two egress manifests, including the sdsmint variant, now use `ALL`, which returns both families and enables Happy Eyeballs. This is a prerequisite for IPv6 egress, not the fix on its own. An actor's connection is redirected by nftables and atunnel recovers the original destination with `getsockopt(SOL_IP, SO_ORIGINAL_DST)`, which returns ENOENT for a v6-redirected connection — so egress fails before Envoy is ever asked to resolve anything. ## Testing `make verify` is clean. A new test walks the shipped manifests and requires `ALL` on every `dns_cache_config`, so a new egress variant cannot reintroduce the pin. It has already paid for itself: the seventh site arrived with the sdsmint MITM leg while this was in review, still pinned to `V4_ONLY`, and the test caught it on rebase. Measured on an IPv6-only kind cluster with agent-substrate#911, agent-substrate#958, agent-substrate#979 and agent-substrate#753 applied. Each value was deployed, Envoy restarted, and the live `/config_dump` checked before running the suites: | `dns_lookup_family` | `TestActorEgress` / `TestActorEgressHTTPS` | | --- | --- | | `ALL` (shipped) | pass / pass, `code=200` to `[2606:4700:10::ac42:93f3]:80` | | `V4_ONLY` (before) | fail / fail, `code=503 flags=DF` | | `AUTO` | pass / pass | `AUTO` is not distinguishable from `ALL` on this path: the dynamic forward proxy is handed the IP literal that atunnel recovered, never a hostname, so `dns_lookup_family` only decides which literal families it accepts. `ALL` is chosen as the value that strands neither family. The IPv4 lane is unaffected: `TestActorEgress` and `TestActorEgressHTTPS` pass with the change in place, across two runs of the standard e2e job. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR 🤖 Generated with [Claude Code](https://claude.com/claude-code)
…ay (agent-substrate#1328) Envoy applies a route timeout to a CONNECT tunnel's whole lifetime, so the 15s default capped every actor's outbound connection. The sdsmint variant already disabled it; add the same line here, and a test over both manifests so they cannot drift apart again. Fixes #<issue_number_goes_here> > It's a good idea to open an issue first for discussion. - [ ] Tests pass - [ ] Appropriate changes to documentation are included in the PR
…ake fix] (agent-substrate#1320) ate-controller creates a WorkerPool's Deployment, so it does not exist yet when the apply returns. `kubectl rollout status` reads the object before it starts watching and errors on a missing one instead of waiting, so whether these waits worked at all came down to beating the controller by a round trip. e2e lost that race by 178ms, failing the deploy with NotFound while the pool's pods were already being created. Fixes yet another flake I observed here: https://github.com/agent-substrate/substrate/actions/runs/33275880827/job/99162382080?pr=1283
…-substrate#1318) Installing the control plane pulls roughly 570MB of third-party images (postgres, prometheus, the otel collector, rustfs, envoy, jaeger) onto a single node. kubelet serializes image pulls by default, so they queue behind one another and whichever workload lands at the back of the queue can miss its readiness deadline. e2e has failed repeatedly with postgres still in PodInitializing after 60s, its image not yet pulled, while everything else was already running. Sample failure: https://github.com/agent-substrate/substrate/actions/runs/33258860897/job/99117220741 Note: This is after PR merge, and it's clearly unrelated to the PR. AI-assisted.
…gent-substrate#1316) Guarding on `[ ! -d "$VENV_DIR" ]` treats an interrupted create (no bin/activate) and an interpreter upgrade (dangling bin/python3) as a good venv, so the script dies in `source venv/bin/activate` and the only fix is knowing to delete the directory. Probe that the venv runs, and rebuild with --clear to relink the interpreter. Also install requirements unconditionally, which is cheap. The license verifier skipped the install whenever the venv already existed, so it passed without ever seeing a newly added dependency. Fixes failure encountered by @ahmedtd. Not filing an issue because this is pretty trivial. AI-assisted.
API reviewers should memorize this :) @EItanya @juli4n @BenTheElder @dberkov
We should not make workerImage optional as it provide explicit information for what image workers use, removing it will need `atecontroller` to inject workerImage and leads to unnecessary control plane & dataplane wiring. Fixes agent-substrate#861 > It's a good idea to open an issue first for discussion. - [ ] Tests pass - [ ] Appropriate changes to documentation are included in the PR
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
August 31, 2026 17:47
1cb5a78 to
7d9aa39
Compare
…e and name pair in internal APIs (agent-substrate#1304) Both Atelet and Ateom internal APIs consume actor template namespace/name pair to telemetry. Updated them to use the new Substrate ActorTemplate proto instead of the legacy CRD one.
The bundle rootfs comes from an untrusted actor image, and the plain os.MkdirAll/os.WriteFile pair followed symlinks out of it: an image with /etc or /etc/resolv.conf pointing at a worker pod path had ateom, running as root, truncate and overwrite that path. Confine the write to the rootfs with os.Root and unlink any existing entry instead of truncating through it.
The proc/sys/dev mountpoints were created with os.MkdirAll on paths inside the container rootfs, whose lower is the actor's image and whose upper is restored from the guest's own snapshot. A symlink planted at one of those names sent the mkdir wherever it pointed on the worker pod, created by ateom as root. Confine it to the rootfs with os.Root, in both the merged-overlay and the reconstructed-from-image staging paths.
* Renames the`ate.template.namespace` attribute to `ate.template.atespace` and sources it from the actor's `actor_template` ObjectRef instead of the legacy CRD namespace/name pair. * Updated functional tests now to use substrate ActorTemplates and bind actors through the `actor_template` ref.
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
August 31, 2026 21:31
7d9aa39 to
3c66a5d
Compare
Fixes agent-substrate#1066 - [x] Tests pass - [x] ~~Appropriate changes to documentation are included in the PR~~ No documentation changes needed. ## Summary - Move actor-facing Envoy upstream mTLS credentials from inline file `DataSource`s to SDS secrets served over ADS. - Publish separate SDS secrets for the upstream podidentity client cert and upstream trust bundle. - Preserve SPIFFE URI SAN validation by using a combined validation context: trust bundle via SDS, SAN matcher inline. - Reuse the ADS config source helper for downstream TLS SDS references. - Add tests covering upstream SDS references, snapshot secret publication, and projected-volume rotation behavior. --------- Signed-off-by: Sachin Kamboj <skamboj1@bloomberg.net>
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
August 31, 2026 22:14
3c66a5d to
3d8fa87
Compare
Before, on an IPv6-only cluster the script rewrote kind's `https://[::1]:PORT` kubeconfig entry to `https://localhost:PORT` unconditionally. That breaks any host whose `/etc/hosts` leaves `localhost` off the `::1` line, including the Ubuntu cloud image Lima runs. Remove the repoint logic, but note that limactl re-forwards the published port to the host's v4 loopback only, so a macOS client cannot directly access kind clusters run inside a Lima VM via `[::1]`. Instead, operations should run inside the guest. Tested: manual tests creating IPv6-only kind clusters on a local env with IPv6 egress, namely macOS + Lima VM.
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
September 3, 2026 20:22
3d8fa87 to
a447f37
Compare
CoreDNS inherits the node's IPv4 resolver, which a v6-only pod cannot reach. Change CoreDNS's Corefile to forward to an overridable IPv6 upstream. Also add a kind-registry:53 server block to CoreDNS. In addition, `kind create` returns before the apiserver answers, so add logic to wait for the control plane from inside the node first. Tested: manual tests of creation and installation on a local env with IPv6 egress, namely macOS + Lima VM. CI cannot verify this — the GitHub runner is IPv4-only, so an IPv6-only cluster there additionally needs DNS64 and NAT64.
ygao-g
force-pushed
the
kind-ipv6-coredns
branch
from
September 3, 2026 20:26
a447f37 to
6c9d0ce
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Review mirror of upstream agent-substrate#958 — do not merge.
Replaces #9, which GitHub will not reopen: its recorded head was force-pushed
away by the rebase onto current
main.The base is pinned to
kind-ipv6-coredns-base(f43d4096), the commit agent-substrate#958 isrebased onto, so the diff here stays exactly the two commits under review even
as upstream
mainmoves. Head tracks the upstream PR branch; leave inlinecomments here and they will be folded back into agent-substrate#958.
01693d31[::1]failsf3e7447aThe NAT64 and DNS64 setup that briefly sat on this branch has moved to its own
upstream pull request, agent-substrate#1275, so what is left here is the part every IPv6-only
cluster needs regardless of whether the host can reach IPv4-only destinations.
🤖 Generated with Claude Code