Skip to content

hack: #958 mirrored for review — fix DNS on IPv6-only kind clusters - #12

Closed
ygao-g wants to merge 58 commits into
kind-ipv6-coredns-basefrom
kind-ipv6-coredns
Closed

ygao-g wants to merge 58 commits into
kind-ipv6-coredns-basefrom
kind-ipv6-coredns

Conversation

@ygao-g

@ygao-g ygao-g commented Aug 27, 2026 •

Copy link
Copy Markdown
Owner

Review mirror of upstream agent-substrate#958 — do not merge.

Replaces #9, which GitHub will not reopen: its recorded head was force-pushed
away by the rebase onto current main.

The base is pinned to kind-ipv6-coredns-base (f43d4096), the commit agent-substrate#958 is
rebased onto, so the diff here stays exactly the two commits under review even
as upstream main moves. Head tracks the upstream PR branch; leave inline
comments here and they will be folded back into agent-substrate#958.

commit subject
01693d31 hack: repoint the IPv6 kubeconfig only when [::1] fails
f3e7447a hack: fix DNS on IPv6-only kind clusters

The NAT64 and DNS64 setup that briefly sat on this branch has moved to its own
upstream pull request, agent-substrate#1275, so what is left here is the part every IPv6-only
cluster needs regardless of whether the host can reach IPv4-only destinations.

🤖 Generated with Claude Code

Comment thread hack/create-kind-cluster.sh Outdated
Comment on lines +265 to +271
# Plain DNS64 synthesizes only for names with no AAAA, and the names that
# matter here have real AAAA records pointing at addresses a host with no
# IPv6 egress cannot reach. Only translate_all forces them through the
# prefix -- which is why this is opt-in: where IPv6 egress does work it
# replaces reachable answers with unreachable ones.
#
# translate_all cannot share a server block with the cluster zones. dns64

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Move these comments to before the corresponding logic.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this applies from 265 to 283

@ygao-g
ygao-g force-pushed the kind-ipv6-coredns branch from a22c540 to f3e7447 Compare August 27, 2026 20:28
@ygao-g ygao-g changed the title hack: #958 mirrored for review — make IPv6-only kind clusters usable hack: #958 mirrored for review — fix DNS on IPv6-only kind clusters Aug 27, 2026
bowei and others added 4 commits August 27, 2026 15:06
Adds ate-setup, a Go-implemented version of the hack scripts
    
    This was done with heavily machine-translated code iterated
    with testing.
…ds (agent-substrate#1276)

The Capabilities, SecurityContext, ImageVolumeSource and Volume.image
fields landed without doc comments, so apitool's documented rule fails
on them and they are not in the exemption backlog. Document them rather
than growing the backlog.

Fixes failure on main

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
agent-substrate#232

Add documentation guide for CSI volumes in Substrate


- [X ] Tests pass
- [ X] Appropriate changes to documentation are included in the PR
…gent-substrate#1263)

The counter tutorial's headline claim — state preserved across
suspend/resume — is false on the default path: with `onCommit: Data`,
suspend excludes process memory, so the in-memory counter resets while
only the durable file counter continues.

**Reproduced live** (GKE, template config identical to main): memory 5 →
**1**,2,3 after suspend/resume; suspend manifest `scope:"data"`,
`pages.img` 71 B (vs 2.2 MB golden Full). The micro-VM variant has the
same broken promise via a different path: `fromData: Golden` restarts
the count from the golden state — the e2e Golden case itself asserts
memory=1 after suspend.

**Fix** (per maintainer decision — config *and* docs, both variants
matching): `onCommit: Full` on both counter templates, drop the now-dead
`onResume` block on the micro-VM one, and make the README precise about
which scope preserves what.

**Verified live**: with `Full`, memory count runs 5 → **6**,7,8 across
suspend/resume; snapshot manifest `scope:"full"`, `pages.img` 2.4 MB.

No coverage lost: `TestDurableDirLifecycle` exercises every
`onCommit`/`onPause`/`fromData` combination (including Data and Golden)
with its own parameterized templates, independent of the demo defaults —
and its `Full/Full` case already asserts the memory count continues, in
both sandbox classes.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
@ygao-g
ygao-g force-pushed the kind-ipv6-coredns branch from f3e7447 to c711eca Compare August 27, 2026 23:31
zoez7 and others added 7 commits August 27, 2026 19:35
…bstrate#1264)

Create/Suspend/Resume Actor workflows now uses the
`Actor.actor_template` ObjectRef when specified, otherwise falls back to
the CRD in cmd/ateapi/internal/controlapi/template_convert.go.
atelet's spec was written for runsc: a hardcoded `runsc` hostname, CRI
pause/sandbox annotations and `dev.gvisor.spec.mount.*` hints. The
micro-VM shim compensated by rewriting every spec at runtime
(ensureKataCompatibleSpec), swapping the whole mount set and rebuilding
it from one hand-written builder per volume kind — so a volume kind
atelet learned about reached the guest only once the shim was taught
about it too.

atelet now emits a spec that names no runtime, and each ateom shapes it
for the runtime it drives. Whatever atelet adds therefore reaches both,
and a bind the micro-VM shaper cannot place in the guest is an error
rather than a silent drop.

Fixes agent-substrate#709.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
…rkloads (--mem-targetLarge mem bench (agent-substrate#1130)

## What this PR does:

Adds large-memory benchmark suites: glutton actors that hold a resident
1–2Gi working set, so the suspend/resume path is measured at realistic
application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm},
tracked.

The boomer GluttonUser fills each actor to a configured target via
chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size
field is int32) after resume and **before the first suspend**, so every
snapshot from cycle one onward carries the full working set. `WriteRAM`
writes incompressible random bytes, so snapshots are genuinely
target-sized rather than zstd-compressing away (a zero-filled working
set compresses ~35:1 and would shrink the upload/download phases to a
few MB).

## How it's configured

The target flows through the established boomer dynconfig channel, like
the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` →
`/boomer-config` → the Go worker. Because it's per-user-class runtime
config, heterogeneous and changing workload shapes are expressible with
no redeploy, and the deploy stack needs no changes.

Changes:
- `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it
was gRPC-only, unreachable in `--mode=http` deployments)
- `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user
cycle — runs once per actor (glutton holds the allocations across
suspend/resume), retried on failure, reported as its own
`GluttonFillRAM` stats row so it never pollutes ping/resume numbers
- `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`:
new `mem_target_bytes` knob
- `tests.yaml`: the four tracked suites, with `actorMemory` sized above
the target for headroom

## Behavior notes

- The golden template snapshot and each actor's cold boot stay small:
the working set exists from the first fill onward. All steady-state
suspend/resume cycles measure at size; `ResumeActorColdStart` does not.
- Each actor's first suspend uploads the first at-size snapshot and
follows the fill — cycle one is an expected outlier on the dashboards.
- A mid-run change to `mem_target_bytes` applies to newly spawned users;
already-filled actors keep their size, so each actor's cycles stay
comparable.
- Existing suites are untouched: with no flag the target is 0 and the
fill is a no-op.

## Testing

- `go test -race` across `cmd/benchmarking/glutton`,
`internal/benchmarking/boomer/...`, and the glutton fake: PASS. New
tests cover chunking to an exact target, disabled-by-default, and
fail-then-retry.
- Cluster verification (microvm, 1Gi target): verified that the RAM fill
completes before the initial suspend, snapshot upload reflects the
expected ~1Gi payload, and steady-state suspend/resume cycles succeed
without error.
…t-substrate#1279)

agent-substrate@4809097#diff-51a595730fdeea0b1efc0394d1b3e866426a9acb000b6f2110262867b9aaeae7
added exemptions unrelated to the PR

agent-substrate#1276 removed the need
for exemptions, which were currently failing on main

This PR cleans up the now failing defunct exemptions.
That's twice today that we missed regenerating the proto clients:

1. mid-PR on agent-substrate#1276
(caught by code review)
2. agent-substrate#1130 (missed and
still missing)

So this PR does:

1. regenerate with latest after agent-substrate#1130
2. ensure verify will catch this (so CI should fail if we haven't run
it), and that update will run it when developing

fixes agent-substrate#738

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
On a fresh GCP project nothing documented which IAM bindings atelet
needs: setup-gcp creates them silently, and anyone who cannot run the
tool with project-level IAM permissions - or needs to audit what it
did - had to read cmd/iam.go and cmd/bucket.go. Spell out the exact
members, roles, and resources, the Workload Identity prerequisites,
and the gcloud equivalents, and point the README quickstart at it.
certificates.k8s.io/v1beta1 podcertificaterequests and
clustertrustbundles cannot be enabled in place on an existing cluster -
the update is accepted but the APIs never become served, and the
install hangs waiting for ClusterTrustBundles. Warn in the create
cluster docs, show the bring-your-own-cluster flag, and name the
symptom.
@ygao-g
ygao-g force-pushed the kind-ipv6-coredns branch from c711eca to 1cb5a78 Compare August 28, 2026 16:23
juli4n and others added 13 commits August 28, 2026 10:09
…#1286)

When checking documentation, remove `validation-gen` tags so we don't
have false negative results if a field *only* has tags and no
documentation.
…ate resource (agent-substrate#1281)

* Added kubectl ate client changes for applying the counter demo.
* Added demo in both gVisor and microVM.
* Updated documentation.

Example output:

```
$ kubectl ate get actortemplates -a ate-demo-counter-substrate
ATESPACE                     NAME      SANDBOX CLASS          STATUS   AGE
ate-demo-counter-substrate   counter   SANDBOX_CLASS_GVISOR   Ready    43s

$ kubectl ate get actortemplate counter -a ate-demo-counter-substrate -o yaml
actorTemplates:
- containers:
  - command:
    - /ko-app/counter
    - --extra-port=9090
    - --tcp-port=9091
    image: gcr.io/zoezhao-gke-dev/ate-images/counter-7b2c368808ac33f45c7ab87955715526@sha256:9eb06b137bfd9f74f7ffa7c6052bb189a7b448801da81931a2124600a2622fa5
    name: counter
    readyz:
      httpGet:
        path: /readyz
        port: 80
    volumeMounts:
    - mountPath: /home/counter
      name: data
  metadata:
    atespace: ate-demo-counter-substrate
    createTime: "2026-08-28T02:09:45.981177255Z"
    name: counter
    uid: a8520ce2-f1f5-45d9-8d94-a436a72c5bf1
    updateTime: "2026-08-28T02:09:56.512455175Z"
    version: "3"
  resources:
    limits:
    - name: cpu
      quantity: "1"
    - name: memory
      quantity: 512Mi
  sandboxConfig:
    configName: gvisor-default
    sandboxClass: SANDBOX_CLASS_GVISOR
  snapshotsConfig:
    onCommit: SNAPSHOT_CONTENT_SCOPE_FULL
    onPause: SNAPSHOT_CONTENT_SCOPE_FULL
    storageLocation: gs://snapshot-substrate-test-zoezhao-gke-dev/ate-demo-counter-substrate/
  status:
    goldenSnapshotStatus:
      goldenSnapshot:
        atespace: ate-demo-counter-substrate
        name: 208d6f11-92e7-4ac0-9cab-74cd0b8a7992
      takeGoldenSnapshotAt: "2026-08-28T02:09:56.041417048Z"
  volumes:
  - durableDir: {}
    name: data
    type: DurableDir
  workerSelector:
    matchLabels:
      workload: counter-substrate
```

More testing log in gpaste/5520785792434176
Flesh out the agentgateway implementation as a router and egress PEP.
This also changes the atunnel CONNECT port to do 8443. Most of this is
agentgateway contained, but the one other system we touch is localca;
agentgateway's MITM support requires the TLS cert and key in the secret
instead of the JSON pool (we don't rely on sdsmint).

- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
…n) (agent-substrate#1280)

#### What this pr do

Re-dirties part of the glutton working set on every benchmark cycle, so
repeated suspends snapshot an actor whose memory is changing — like a
live application's — instead of a set that was filled once and never
touched again.

Each iteration (after the one-time fill), the GluttonUser sends one
`WriteRAM` request with `WRITE_MODE_OVERWRITE`, re-randomizing the first
`mem_churn` bytes of the working set in place. Overwrite mode already
existed on the glutton API; this PR adds no proto changes and no glutton
server changes — it is driver + config wiring only.

This is the follow-up from agent-substrate#1130's: the original self-driving memload
continuously re-dirtied its pages, and that property was dropped in the
move to API-driven fill. This restores it in the API-driven shape with a
dialable amount instead of the old all-or-nothing full pass.

#### Why it matters

A fill-once working set is static: every suspend after the first packs
up identical memory. If snapshotting ever gains incremental / dirty-page
optimizations, a static benchmark would measure almost nothing from
cycle two onward — and would score "upload nothing" as an infinite win.
With churn, every cycle carries a known amount of freshly-dirtied pages,
and the knob is sweepable (64Mi, 256Mi, …) so a future
differential-snapshot optimization can be demonstrated as "upload cost
scales with churn size, not total size."
…strate#1210)

Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.

This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.

Part of agent-substrate#861

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
…strate#1285)

The networking suite's direct-access and arbitrary-port tests build
their actors from the counter fixture and behave the same whichever
egress gateway is deployed, so re-running them after the sdsmint swap
adds CI time without adding coverage. Only egressFixture() reads
E2E_EGRESS_MITM, so restrict both MITM lanes to the tests that use it.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
BenTheElder and others added 17 commits August 28, 2026 21:08
Necessary for agent-substrate#1266, but functionally useful regardless.

This eliminates the MPL go-reap dependency in favor of our own
implementation tuned to our needs.

Our package has utilities for spawning child processes without holding a
RWMutex to avoid convoying lots of concurrent activations in the future.
Add maxItems checks to lists and simplify uniqueness checks.
…bstrate#595)

DNS lookups are case-insensitive (RFC 4343), so a client that resolves
`MyActor.MySpace.actors.resources.substrate.ate.dev` reaches the router
with that spelling preserved in `Host`/`:authority` (Go's http.Client
and `curl -H` both send the host as written). `ParseActorDNSName`
compared it byte for byte against a strictly lower-case suffix and name
pattern, so the request was rejected with a 404 by `invalidHostErr` even
though its DNS lookup had resolved to that very actor.

Fold the name to lower case before parsing. Actor and atespace names are
always lower case (`ResourceNameRegexPattern`), so a folded name either
addresses the same actor its lookup resolved to, or fails validation
exactly as before.
…#1300)

Remove `ActorTemplateStatus.sandbox_assets` along with the
`SandboxAssets` messages it referenced. Nothing ever wrote the field:
sandbox assets are resolved from the WorkerPool and SandboxConfig
objects at resume time and travel to atelet via ateletpb, so freezing
them into the template status never materialized.

This requires more thought, one option is to add it as a field of
GoldenSnapshotStatus in the future, and make GoldenSnapshotStatus a
repeated field to allow multiple goldens.
…#416)

Part of agent-substrate#207, following the plan discussed there with @juli4n and
@kannon92: plumbing plus the non invasive fixes first, with the existing
invasive findings excluded so new APIs do not regress (as @kannon92
suggested in
agent-substrate#207 (comment)).
The linter itself was originally suggested by @BenTheElder in agent-substrate#188.

Plumbing:
- hack/tools/kube-api-linter module with the
golangci-lint-kube-api-linter tool, same pattern as the other tools.
- .golangci-kal.yaml config. The pre-existing findings from rules that
need Go API changes (nomaps, nonpointerstructs, nophase, optionalfields,
requiredfields, ssatags) are excluded rather than disabled, and each
exclusion rule names the fields it excuses, so the rules still apply to
new API files and to new fields on the existing types. Fixing them on
the existing types is the follow up in agent-substrate#207.
- hack/verify/kube-api-linter.sh runs it against ./pkg/api/... and is
picked up by hack/verify-all.sh in CI.

Fixes in pkg/api/v1alpha1 (doc and marker only, no Go API change):
- commentstart: field godocs now start with the serialized field name.
- conditions: Conditions moved to the first position in
ActorTemplateStatus and got the listType, listMapKey and patch markers.
- defaultorrequired: SandboxConfigSpec.SandboxClass had both a default
and required; now optional with omitempty, matching the same field on
WorkerPoolSpec and ActorTemplateSpec. The ValidatingAdmissionPolicy in
sandboxconfig-validation.yaml is unaffected since CRD defaulting runs
before admission, so spec.sandboxClass is always set by then.
- optionalorrequired: missing +optional markers added on
ActorTemplateStatus fields.
- defaults: configured preferredDefaultMarker to kubebuilder:default
since CRDs are generated with controller-gen.

Generated CRDs regenerated. Structural schema changes are only the
conditions listType/listMapKey and sandboxClass no longer in the
required list (it is defaulted, so behavior is the same); the rest is
description text.

Note: my earlier count of 50 issues in agent-substrate#207 was capped by the default
golangci-lint issue limit. With the cap removed the real total is 86;
the extra ones are requiredfields (20) and ssatags (4), both in the
excluded set above.

Test plan
- hack/verify/kube-api-linter.sh passes (0 issues).
- Exclusions verified both ways: dropping them brings the pre-existing
findings back, and a scratch field added to WorkerPoolSpec is still
reported (optionalfields), which the earlier per-file exclusion
swallowed.
- TestSandboxConfigValidation passes against the regenerated CRD.
- go build ./... and go test ./pkg/api/... ./cmd/atecontroller/... pass.
- gofmt, shellcheck and the regular golangci-lint on pkg/api are clean.
ateapi is a headless Service, so its DNS name resolves to one address
per replica and the client picks. This one asked for no policy, which
means gRPC's pick_first: it connects to whichever address answers first
and sends everything there for the life of the connection.

So a client talks to one replica however many are running, and adding
replicas moves no load. Measured with eight load generators against two
replicas, one took 97% of the traffic.

internal/ateapiauth, which atelet and the controllers dial through, has
asked for round_robin all along. This is the same policy for the clients
that do not.

Pulled out from agent-substrate#1266, I ran into this while doing some custom
benchmarking using ateclient, but this is already a bug for kubectl-ate,
demos, ...

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

AI-assisted.
…-substrate#1060)

Part of agent-substrate#246

The egress Envoy pinned `dns_lookup_family` to `V4_ONLY` in both the
dynamic
forward proxy filter and the dynamic forward proxy cluster. On an
IPv6-only
cluster every upstream connection failed: Envoy reported a DNS
resolution
failure and the actor got a 503. Seven sites across the two egress
manifests,
including the sdsmint variant, now use `ALL`, which returns both
families and
enables Happy Eyeballs.

This is a prerequisite for IPv6 egress, not the fix on its own. An
actor's
connection is redirected by nftables and atunnel recovers the original
destination with `getsockopt(SOL_IP, SO_ORIGINAL_DST)`, which returns
ENOENT
for a v6-redirected connection — so egress fails before Envoy is ever
asked to
resolve anything.

## Testing

`make verify` is clean. A new test walks the shipped manifests and
requires
`ALL` on every `dns_cache_config`, so a new egress variant cannot
reintroduce
the pin. It has already paid for itself: the seventh site arrived with
the
sdsmint MITM leg while this was in review, still pinned to `V4_ONLY`,
and the
test caught it on rebase.

Measured on an IPv6-only kind cluster with agent-substrate#911, agent-substrate#958, agent-substrate#979 and agent-substrate#753
applied.
Each value was deployed, Envoy restarted, and the live `/config_dump`
checked
before running the suites:

| `dns_lookup_family` | `TestActorEgress` / `TestActorEgressHTTPS` |
| --- | --- |
| `ALL` (shipped) | pass / pass, `code=200` to
`[2606:4700:10::ac42:93f3]:80` |
| `V4_ONLY` (before) | fail / fail, `code=503 flags=DF` |
| `AUTO` | pass / pass |

`AUTO` is not distinguishable from `ALL` on this path: the dynamic
forward
proxy is handed the IP literal that atunnel recovered, never a hostname,
so
`dns_lookup_family` only decides which literal families it accepts.
`ALL` is
chosen as the value that strands neither family.

The IPv4 lane is unaffected: `TestActorEgress` and
`TestActorEgressHTTPS` pass
with the change in place, across two runs of the standard e2e job.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

🤖 Generated with [Claude Code](https://claude.com/claude-code)
…ay (agent-substrate#1328)

Envoy applies a route timeout to a CONNECT tunnel's whole lifetime, so
the 15s default capped every actor's outbound connection. The sdsmint
variant already disabled it; add the same line here, and a test over
both manifests so they cannot drift apart again.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
…ake fix] (agent-substrate#1320)

ate-controller creates a WorkerPool's Deployment, so it does not exist
yet when the apply returns.

`kubectl rollout status` reads the object before it starts watching and
errors on a missing one instead of waiting, so whether these waits
worked at all came down to beating the controller by a round trip.

e2e lost that race by 178ms, failing the deploy with NotFound while the
pool's pods were already being created.

Fixes yet another flake I observed here:
https://github.com/agent-substrate/substrate/actions/runs/33275880827/job/99162382080?pr=1283
…-substrate#1318)

Installing the control plane pulls roughly 570MB of third-party images
(postgres, prometheus, the otel collector, rustfs, envoy, jaeger) onto a
single node.

kubelet serializes image pulls by default, so they queue behind one
another and whichever workload lands at the back of the queue can miss
its readiness deadline.

e2e has failed repeatedly with postgres still in PodInitializing after
60s, its image not yet pulled, while everything else was already
running.

Sample failure:
https://github.com/agent-substrate/substrate/actions/runs/33258860897/job/99117220741

Note: This is after PR merge, and it's clearly unrelated to the PR.

AI-assisted.
…gent-substrate#1316)

Guarding on `[ ! -d "$VENV_DIR" ]` treats an interrupted create (no
bin/activate) and an interpreter upgrade (dangling bin/python3) as a
good venv, so the script dies in `source venv/bin/activate` and the only
fix is knowing to delete the directory. Probe that the venv runs, and
rebuild with --clear to relink the interpreter.

Also install requirements unconditionally, which is cheap. The license
verifier skipped the install whenever the venv already existed, so it
passed without ever seeing a newly added dependency.

Fixes failure encountered by @ahmedtd. Not filing an issue because this
is pretty trivial.

AI-assisted.
We should not make workerImage optional as it provide explicit
information for what image workers use, removing it will need
`atecontroller` to inject workerImage and leads to unnecessary control
plane & dataplane wiring.



Fixes agent-substrate#861 

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
@ygao-g
ygao-g force-pushed the kind-ipv6-coredns branch from 1cb5a78 to 7d9aa39 Compare August 31, 2026 17:47
zoez7 and others added 4 commits August 31, 2026 14:50
…e and name pair in internal APIs (agent-substrate#1304)

Both Atelet and Ateom internal APIs consume actor template
namespace/name pair to telemetry. Updated them to use the new Substrate
ActorTemplate proto instead of the legacy CRD one.
The bundle rootfs comes from an untrusted actor image, and the plain
os.MkdirAll/os.WriteFile pair followed symlinks out of it: an image with
/etc or /etc/resolv.conf pointing at a worker pod path had ateom, running
as root, truncate and overwrite that path. Confine the write to the
rootfs with os.Root and unlink any existing entry instead of truncating
through it.
The proc/sys/dev mountpoints were created with os.MkdirAll on paths
inside the container rootfs, whose lower is the actor's image and whose
upper is restored from the guest's own snapshot. A symlink planted at
one of those names sent the mkdir wherever it pointed on the worker pod,
created by ateom as root. Confine it to the rootfs with os.Root, in both
the merged-overlay and the reconstructed-from-image staging paths.
* Renames the`ate.template.namespace` attribute to
`ate.template.atespace` and sources it from the actor's `actor_template`
ObjectRef instead of the legacy CRD namespace/name pair.
* Updated functional tests now to use substrate ActorTemplates and bind
actors through the `actor_template` ref.
@ygao-g
ygao-g force-pushed the kind-ipv6-coredns branch from 7d9aa39 to 3c66a5d Compare August 31, 2026 21:31
Fixes agent-substrate#1066 

- [x] Tests pass
- [x] ~~Appropriate changes to documentation are included in the PR~~ No
documentation changes needed.

## Summary

- Move actor-facing Envoy upstream mTLS credentials from inline file
`DataSource`s to SDS secrets served over ADS.
- Publish separate SDS secrets for the upstream podidentity client cert
and upstream trust bundle.
- Preserve SPIFFE URI SAN validation by using a combined validation
context: trust bundle via SDS, SAN matcher inline.
- Reuse the ADS config source helper for downstream TLS SDS references.
- Add tests covering upstream SDS references, snapshot secret
publication, and projected-volume rotation behavior.

---------

Signed-off-by: Sachin Kamboj <skamboj1@bloomberg.net>
@ygao-g
ygao-g force-pushed the kind-ipv6-coredns branch from 3c66a5d to 3d8fa87 Compare August 31, 2026 22:14
Before, on an IPv6-only cluster the script rewrote kind's
`https://[::1]:PORT` kubeconfig entry to `https://localhost:PORT`
unconditionally. That breaks any host whose `/etc/hosts` leaves
`localhost` off the `::1` line, including the Ubuntu cloud image Lima
runs. Remove the repoint logic, but note that limactl re-forwards the
published port to the host's v4 loopback only, so a macOS client cannot
directly access kind clusters run inside a Lima VM via `[::1]`. Instead,
operations should run inside the guest.

Tested: manual tests creating IPv6-only kind clusters on a local env
with IPv6 egress, namely macOS + Lima VM.
CoreDNS inherits the node's IPv4 resolver, which a v6-only pod cannot
reach. Change CoreDNS's Corefile to forward to an overridable IPv6
upstream. Also add a kind-registry:53 server block to CoreDNS. In
addition, `kind create` returns before the apiserver answers, so add
logic to wait for the control plane from inside the node first.

Tested: manual tests of creation and installation on a local env with
IPv6 egress, namely macOS + Lima VM. CI cannot verify this — the GitHub
runner is IPv4-only, so an IPv6-only cluster there additionally needs
DNS64 and NAT64.
@ygao-g ygao-g closed this Sep 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.