diff --git a/docs/backends/bwrap/bubblewrap-backend.md b/docs/backends/bwrap/bubblewrap-backend.md index 02d9584e1..13afce91d 100644 --- a/docs/backends/bwrap/bubblewrap-backend.md +++ b/docs/backends/bwrap/bubblewrap-backend.md @@ -172,14 +172,17 @@ Common consequences of this default: `readonlyPaths` if the script depends on it. - `working_directory` must live under the baseline or a policy path — a `cwd` of `~/project` without a matching `readonlyPaths` entry will fail. -- DNS works on systemd-resolved, NetworkManager, and resolvconf hosts - because the corresponding `/run/...` directories are bound. The common - symlink targets *outside* `/run` are covered too: `/var/run/...`-routed - `/etc/resolv.conf` symlinks resolve via a synthesised `/var/run -> /run` - compat symlink, and WSL's `/mnt/wsl/resolv.conf` is bound directly. - Neither exposes host `/var` or `/mnt` contents. Hosts that point - `/etc/resolv.conf` at some other custom location still need that target - listed in `readonlyPaths`. +- The sandbox can *read* the host's resolver on systemd-resolved, + NetworkManager, and resolvconf hosts because the corresponding `/run/...` + directories are bound. The common symlink targets *outside* `/run` are + covered too: `/var/run/...`-routed `/etc/resolv.conf` symlinks resolve via a + synthesised `/var/run -> /run` compat symlink, and WSL's + `/mnt/wsl/resolv.conf` is bound directly. Neither exposes host `/var` or + `/mnt` contents. Hosts that point `/etc/resolv.conf` at some other custom + location still need that target listed in `readonlyPaths`. + + Reading the file is not the same as reaching the nameserver it names — see + [Loopback resolvers](#loopback-resolvers). Files in `/etc` that contain secrets (`/etc/shadow`, `/etc/sudoers`, `/etc/ssh/ssh_host_*_key`) are mode `0400` / `0640` `root` and remain @@ -302,6 +305,59 @@ does not resolve, because the sandbox resolves names itself and a lookup that disagreed with the one behind the rules would hand the workload an address the chain never authorized. +### Loopback resolvers + +Every mode runs the sandbox in its own network namespace, so a nameserver on +the host's loopback — `systemd-resolved`'s stub at `127.0.0.53`, a local +`dnsmasq` at `127.0.0.1` — names an address that belongs to the sandbox's own +empty loopback. Binding the host's resolver directories makes the file +readable; it does not make that address reachable, and names do not resolve. + +Under **address filtering**, MXC therefore replaces the sandbox's resolver with +slirp's built-in forwarder, `10.0.2.3`. libslirp rewrites a query sent there to +the first IPv4 nameserver in the *host's* `/etc/resolv.conf` and sends it from +the host's network namespace, which is what puts the host's stub back in reach. +`search`, `options`, and every other directive are carried over unchanged. + +The replacement applies only when the host's own resolvers are **all** loopback +addresses and **at least one** is IPv4, so it cannot redirect name resolution +that already works, and it never applies when slirp has no IPv4 nameserver to +forward to. The file is mounted over the path the `/etc/resolv.conf` symlink +chain ends at — resolved through every component, not just the leaf, because +bwrap cannot create a mount point under an unresolved symlink and aborts the +sandbox when asked to. + +**The pin never costs you a sandbox.** It improves on a resolver the sandbox +already cannot reach, so every failure gives up the pin rather than the run: +a resolver path that cannot be followed, one the filesystem policy denies, a +resolver file that exists but cannot be read, and a failure to write the +replacement all leave the sandbox exactly as it would have started without +this feature, each with a warning naming the cause. A host with no +`/etc/resolv.conf` at all is silent — there is no loopback stub to rescue. + +A policy that names `/etc/resolv.conf` (or the path its symlink chain ends at) +in `readonlyPaths` / `readwritePaths` is supplying its own resolver, and the +sandbox keeps that one. Listing an ancestor such as `/etc` is not treated as a +choice about the resolver. + +Two cases are deliberately untouched: + +- **`runtimeConfig.networkProxy`** — a proxy-only chain opens no port 53 and + the proxy resolves on the workload's behalf. +- **Ruleless `egress.default: "deny"`** — the sandbox has no connectivity at + all, so it has nothing to resolve with. + +Under `egress.default: "deny"` *with* rules, the chain's terminal verdict still +governs: a query to `10.0.2.3` is dropped unless a rule permits it. MXC warns +at launch when it pins the resolver and the chain admits nothing to +`10.0.2.3:53`, naming the rule to add — allow `10.0.2.3/32` on `udp` port 53. +The rule is not added automatically: the chain is the caller's policy. + +> ⚠️ slirp's forwarder reaches the host's resolver by design, and the chain's +> `-d 10.0.2.2/32 -j DROP` rule — which closes the sandbox's path to host +> loopback services — does not cover it. Name resolution is the one host +> loopback service an address-filtered sandbox can reach. + **IPv6 rules are filtered, but IPv6 traffic has nowhere to go.** The two are separate concerns and only the second is missing. Filtering works: an IPv6 rule programs `ip6tables`, and the terminal verdict of the unmatched family follows diff --git a/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_command.rs b/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_command.rs index 34a023565..135c25bb7 100644 --- a/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_command.rs +++ b/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_command.rs @@ -45,14 +45,20 @@ pub(crate) const COMMAND_TAIL: [&str; 3] = ["--", "sh", "-c"]; /// - We deliberately do NOT bind `/run` wholesale: `/run/user/` /// holds the caller's D-Bus session socket, keyring sockets, and /// ssh-agent socket. We only bind the well-known DNS stub-resolver -/// directories so name resolution still works when `/etc/resolv.conf` +/// directories so the resolver file is *readable* when `/etc/resolv.conf` /// is a symlink (the default on systemd-resolved hosts). -/// - To keep DNS working when `/etc/resolv.conf` points *outside* those -/// dirs, we also synthesise a `/var/run -> /run` compat symlink (for +/// - To keep that file readable when `/etc/resolv.conf` points *outside* +/// those dirs, we also synthesise a `/var/run -> /run` compat symlink (for /// `/var/run/...`-routed targets — older RHEL/CentOS-era and some /// container images) and `--ro-bind-try` `/mnt/wsl/resolv.conf` (for /// WSL). Neither exposes host `/var` or `/mnt` contents — only the /// resolver path itself. +/// - Readable is not reachable. A nameserver on the host's loopback — the +/// systemd-resolved stub at `127.0.0.53`, a local `dnsmasq` — belongs to +/// the sandbox's own empty loopback once it is in a private network +/// namespace, so these binds alone leave names unresolvable. Reaching such +/// a resolver is what `proxy_network::ResolverPin` does, by pointing the +/// sandbox at slirp's forwarder. /// - `/etc` is bound whole because cherry-picking files (`passwd`, /// `nsswitch.conf`, `ssl/`, `ld.so.conf*`, …) is fragile and breaks /// tools that read other config files. Files with sensitive contents @@ -1285,10 +1291,10 @@ mod tests { } } - /// DNS stub-resolver dirs must be in the baseline so `/etc/resolv.conf` - /// symlinks resolve when the caller has network access. Emitted via - /// `--ro-bind-try` so hosts without systemd-resolved / NetworkManager / - /// resolvconf still build a valid argument vector. + /// DNS stub-resolver dirs must be in the baseline so an `/etc/resolv.conf` + /// symlink has a target to resolve to. Emitted via `--ro-bind-try` so hosts + /// without systemd-resolved / NetworkManager / resolvconf still build a + /// valid argument vector. #[test] fn baseline_includes_dns_stub_resolver_dirs() { let args = build_args(&base_request(), None); @@ -1302,8 +1308,8 @@ mod tests { .any(|w| w[0] == "--ro-bind-try" && w[1] == path && w[2] == path); assert!( found, - "baseline must emit `--ro-bind-try {} {}` so DNS works when \ - /etc/resolv.conf is a symlink", + "baseline must emit `--ro-bind-try {} {}` so the sandbox can read \ + /etc/resolv.conf when it is a symlink", path, path ); } diff --git a/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_runner.rs b/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_runner.rs index edbcfaccb..1829df430 100644 --- a/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_runner.rs +++ b/src/mxc-sdk/src/backends/bubblewrap/common/bwrap_runner.rs @@ -17,11 +17,12 @@ use std::collections::HashSet; use std::fmt::Write as FmtWrite; use std::os::fd::AsFd; use std::os::unix::process::CommandExt; -use std::path::{Component, Path, PathBuf}; +use std::path::Path; use std::process::{Child, ChildStdin, Command, ExitStatus, Stdio}; use std::sync::{Arc, Mutex, MutexGuard, PoisonError}; use std::time::{Duration, Instant}; +use crate::mxc_common::filesystem_symlink::resolve_through_symlinks; use crate::mxc_common::interruptible_reader::{wrap_pipe, InterruptibleReader, ReadCanceller}; use crate::mxc_common::logger::Logger; use crate::mxc_common::models::{ExecutionRequest, ScriptResponse}; @@ -42,6 +43,8 @@ use crate::bwrap_common::{ proxy_network, }; +const DNS_PORT: u16 = 53; + /// Bubblewrap sandbox runner. Uses only shared `ContainerPolicy` fields — /// no backend-specific config struct required. #[derive(Default)] @@ -350,6 +353,32 @@ fn warn_unreachable_v6_targets(plan: &network_rules::EgressPlan, logger: &mut Lo )); } +/// Warn that the resolver pin cannot work under the installed chain. +/// +/// The pin points the sandbox at slirp's forwarder, but the forwarder is a +/// destination like any other: a chain that does not admit it drops the query +/// and names still fail to resolve. The chain is the caller's policy, so this +/// warns rather than opening the port on their behalf. +fn warn_resolver_blocked_by_egress( + plan: &network_rules::EgressPlan, + resolver: Option<&proxy_network::ResolverPin>, + logger: &mut Logger, +) { + if resolver.is_none() || plan.admits_udp(proxy_network::SLIRP_DNS_FORWARDER_IP, DNS_PORT) { + return; + } + logger.warning_line(&format!( + "WARNING: Bubblewrap pointed the sandbox at slirp's DNS forwarder {} because the host \ + resolves through a loopback nameserver, but network.egress does not admit UDP port {} \ + to it, so queries are dropped and names will not resolve. Add an allow rule for \ + '{}/32' on udp port {} if the workload needs to resolve names.", + proxy_network::SLIRP_DNS_FORWARDER_IP, + DNS_PORT, + proxy_network::SLIRP_DNS_FORWARDER_IP, + DNS_PORT + )); +} + impl BubblewrapScriptRunner { /// Set up networking and spawn `bwrap`, returning a [`BwrapChild`] wrapped /// by the [`SandboxProcess`] handle. With [`StdioMode::Pipes`] the child's @@ -398,10 +427,15 @@ impl BubblewrapScriptRunner { // the egress rule must open. Some(resolved) => { let egress = resolved.egress(); + + // A proxied chain opens no port 53, so the workload resolves + // through the proxy rather than a nameserver of its own. + let resolver = None; match proxy_network::ProxyNetworkNamespace::start( &egress.plan(), &network_rules::IngressPlan::for_policy(&request.policy), egress.pin(), + resolver, logger, request.script_timeout, ) { @@ -422,10 +456,13 @@ impl BubblewrapScriptRunner { Err(error) => return Err(ScriptResponse::error(&error)), }, }; + let resolver = proxy_network::ResolverPin::for_host(&request.policy, logger); + warn_resolver_blocked_by_egress(&plan, resolver.as_ref(), logger); match proxy_network::ProxyNetworkNamespace::start( &plan, &network_rules::IngressPlan::for_policy(&request.policy), None, + resolver.as_ref(), logger, request.script_timeout, ) { @@ -1120,50 +1157,6 @@ fn is_file_mask_target(path: &str) -> bool { .unwrap_or(false) } -/// Resolve every symlink in `path` (leaf and ancestors) to a real filesystem -/// path, tolerating trailing components that do not exist yet. -/// -/// `std::fs::canonicalize` resolves symlinks at every level but requires the -/// **whole** path to exist. To also cover not-yet-created denied paths under a -/// symlinked ancestor, this walks the components from the root: every existing -/// prefix is canonicalized (following symlinks exactly like the kernel), while -/// `.` and `..` in the not-yet-existent tail are folded lexically. Folding `..` -/// this way is safe because a component that does not exist cannot be a symlink, -/// so the result matches the target the kernel's path resolution would reach. -/// Returns `None` only for an empty path. -/// -/// A naive backward walk that collected `file_name()` silently dropped `..` -/// components (Rust returns `None` for a `..` file name) and reconstructed the -/// wrong target: `/link/missing/../secret` became `/real/missing/secret` -/// instead of `/real/secret`, so the mask landed on a bystander path and the -/// real denied target stayed exposed. -fn resolve_through_symlinks(path: &Path) -> Option { - let mut result = PathBuf::new(); - for component in path.components() { - match component { - Component::Prefix(_) | Component::RootDir => result.push(component.as_os_str()), - Component::CurDir => {} - Component::ParentDir => { - result.pop(); - } - Component::Normal(name) => { - result.push(name); - // Canonicalize the prefix so far so symlinks are followed while - // it still exists; once a component is missing, canonicalize - // fails and the remaining tail is folded lexically above. - if let Ok(real) = std::fs::canonicalize(&result) { - result = real; - } - } - } - } - if result.as_os_str().is_empty() { - None - } else { - Some(result) - } -} - #[cfg(test)] mod tests { use super::*; @@ -1580,6 +1573,104 @@ mod tests { assert!(logger.warnings().is_empty(), "no v6 allow, no warning"); } + /// A pin built the way production builds one. + fn loopback_resolver_pin(dir: &tempfile::TempDir) -> proxy_network::ResolverPin { + let path = dir.path().join("resolv.conf"); + std::fs::write(&path, "nameserver 127.0.0.53\nsearch corp.example\n") + .expect("fixture resolver"); + let mut logger = Logger::new(crate::mxc_common::logger::Mode::Buffer); + proxy_network::ResolverPin::from_resolver( + &path, + &crate::mxc_common::models::ContainerPolicy::default(), + &mut logger, + ) + .expect("a loopback stub is pinned") + } + + /// An allowlist that never names slirp's forwarder leaves the pinned + /// resolver unreachable, which the sandbox cannot report for itself: the + /// resolver file looks correct and the query is simply dropped. + #[test] + fn a_pinned_resolver_the_chain_drops_is_warned_about() { + use crate::mxc_common::models::{NetworkEgressPolicy, NetworkPeer, NetworkRule}; + + let mut req = base_request(); + req.policy.network_egress = Some(NetworkEgressPolicy { + allow: vec![NetworkRule { + to: vec![NetworkPeer { + cidr: "140.82.112.0/20".parse().expect("test CIDR"), + except: vec![], + }], + ports: vec![], + }], + ..Default::default() + }); + let plan = + network_rules::EgressPlan::for_request(&req).expect("an allowlist is enforceable"); + let dir = tempfile::tempdir().unwrap(); + let pin = loopback_resolver_pin(&dir); + + let mut logger = Logger::new(crate::mxc_common::logger::Mode::Buffer); + warn_resolver_blocked_by_egress(&plan, Some(&pin), &mut logger); + let out = logger.warnings().join("\n"); + assert!(out.contains("10.0.2.3"), "must name the forwarder: {out}"); + assert!( + out.contains("udp") && out.contains("53"), + "must name the rule the caller needs: {out}" + ); + assert!( + logger.get_buffer().is_empty(), + "the warning must travel as a retained warning, not as buffer output" + ); + + // A host whose resolver already works is never pinned, so there is + // nothing to warn about however closed the chain is. + let mut unpinned = Logger::new(crate::mxc_common::logger::Mode::Buffer); + warn_resolver_blocked_by_egress(&plan, None, &mut unpinned); + assert!( + unpinned.warnings().is_empty(), + "no pin, no warning: {:?}", + unpinned.warnings() + ); + } + + /// Warning at a caller who already opened the forwarder would send them to + /// change a rule that is already correct. + #[test] + fn a_pinned_resolver_the_chain_admits_is_not_warned_about() { + use crate::mxc_common::models::{ + NetworkEgressPolicy, NetworkPeer, NetworkPort, NetworkProtocol, NetworkRule, + }; + + let mut req = base_request(); + req.policy.network_egress = Some(NetworkEgressPolicy { + allow: vec![NetworkRule { + to: vec![NetworkPeer { + cidr: "10.0.2.3/32".parse().expect("test CIDR"), + except: vec![], + }], + ports: vec![NetworkPort { + protocol: NetworkProtocol::Udp, + port: Some(53), + end_port: None, + }], + }], + ..Default::default() + }); + let plan = + network_rules::EgressPlan::for_request(&req).expect("a udp/53 allow is enforceable"); + let dir = tempfile::tempdir().unwrap(); + let pin = loopback_resolver_pin(&dir); + + let mut logger = Logger::new(crate::mxc_common::logger::Mode::Buffer); + warn_resolver_blocked_by_egress(&plan, Some(&pin), &mut logger); + assert!( + logger.warnings().is_empty(), + "the chain admits DNS, so there is nothing to report: {:?}", + logger.warnings() + ); + } + /// Proxy-only mode rewrites the endpoint to slirp's gateway and opens /// exactly that address, so an endpoint neither step can express has to be /// refused at policy time -- before a proxy is started. diff --git a/src/mxc-sdk/src/backends/bubblewrap/common/network_rules.rs b/src/mxc-sdk/src/backends/bubblewrap/common/network_rules.rs index 9521631a7..5c52bb83a 100644 --- a/src/mxc-sdk/src/backends/bubblewrap/common/network_rules.rs +++ b/src/mxc-sdk/src/backends/bubblewrap/common/network_rules.rs @@ -141,6 +141,35 @@ impl RuleAddress { text: format!("{address}/{prefix}"), } } + + /// Whether this destination covers `address`. + /// + /// The text is either a bare literal or a CIDR block, because a proxy + /// endpoint is rendered without a prefix while a lowered peer keeps one. + /// An unparseable block reports no match, which can only make a diagnostic + /// quieter than the chain it describes -- never more permissive. + fn contains_v4(&self, address: Ipv4Addr) -> bool { + if self.family != RuleFamily::V4 { + return false; + } + let (base, prefix) = match self.text.split_once('/') { + Some((base, prefix)) => match prefix.parse::() { + Ok(prefix) if prefix <= 32 => (base, prefix), + _ => return false, + }, + None => (self.text.as_str(), 32), + }; + let Ok(base) = base.parse::() else { + return false; + }; + // `u32::MAX << 32` is undefined, so the /0 case is masked explicitly. + let mask = if prefix == 0 { + 0 + } else { + u32::MAX << (32 - u32::from(prefix)) + }; + (base.to_bits() & mask) == (address.to_bits() & mask) + } } /// An IPv4-mapped CIDR as its IPv4 equivalent, when it has one. @@ -620,6 +649,27 @@ impl EgressRule { self.verdict.target() ) } + + /// Whether this rule matches a UDP datagram to `destination:port`. + /// + /// An absent protocol or port range matches everything, which is how + /// `iptables` reads a rule carrying no `-p` or `--dport`. + fn matches_udp(&self, destination: Ipv4Addr, port: u16) -> bool { + if self.address.family != RuleFamily::V4 { + return false; + } + if !matches!(self.port.protocol, None | Some("udp")) { + return false; + } + if self + .port + .range + .is_some_and(|(start, end)| port < start || port > end) + { + return false; + } + self.address.contains_v4(destination) + } } /// The complete filtering posture for one sandbox. @@ -801,6 +851,21 @@ impl EgressPlan { } targets } + + /// Whether the installed chain lets a UDP datagram reach `destination:port`. + /// + /// The chain is evaluated first-match, so this walks the rules in the order + /// they are installed and takes the verdict of the first that matches, + /// falling back to the family's terminal verdict. The loopback exemption is + /// not consulted: it matches on outbound interface, which no destination + /// address decides. + pub(crate) fn admits_udp(&self, destination: Ipv4Addr, port: u16) -> bool { + self.rules + .iter() + .find(|rule| rule.matches_udp(destination, port)) + .map_or(self.v4_terminal, |rule| rule.verdict) + == RuleVerdict::Accept + } } /// The inbound posture for one sandbox. @@ -1930,6 +1995,139 @@ mod tests { assert!(mapped.allowed_v6_targets().is_empty()); } + /// The forwarder address the resolver pin writes, which these ask the chain + /// about. + const FORWARDER: Ipv4Addr = Ipv4Addr::new(10, 0, 2, 3); + + /// A ruleless allow default opens the forwarder along with everything else. + #[test] + fn an_open_egress_default_admits_the_dns_forwarder() { + let plan = EgressPlan::for_request(&request(NetworkAction::Allow, &[], &[])) + .expect("an open default is enforceable"); + + assert!(plan.admits_udp(FORWARDER, 53)); + } + + /// The case the warning exists for: an allowlist that never names the + /// forwarder drops the query at the chain's terminal verdict. + #[test] + fn a_deny_default_without_a_dns_rule_blocks_the_forwarder() { + let plan = + EgressPlan::for_request(&request(NetworkAction::Deny, &["140.82.112.0/20"], &[])) + .expect("an allowlist is enforceable"); + + assert!(!plan.admits_udp(FORWARDER, 53)); + } + + /// The remedy the warning names has to actually silence it. + #[test] + fn an_explicit_dns_rule_admits_the_forwarder() { + let req = directional_rules_request( + NetworkAction::Deny, + vec![rule( + vec![peer("10.0.2.3/32", &[])], + vec![port(NetworkProtocol::Udp, Some(53), None)], + )], + Vec::new(), + ); + let plan = EgressPlan::for_request(&req).expect("a udp/53 allow is enforceable"); + + assert!(plan.admits_udp(FORWARDER, 53)); + // The rule names one port, so it must not read as opening others. + assert!(!plan.admits_udp(FORWARDER, 443)); + } + + /// A rule the forwarder merely falls inside counts, so an operator who + /// opened a wider range is not told to add one they already have. + #[test] + fn a_containing_block_admits_the_forwarder() { + let req = directional_rules_request( + NetworkAction::Deny, + vec![rule( + vec![peer("10.0.2.0/24", &[])], + vec![port(NetworkProtocol::Any, None, None)], + )], + Vec::new(), + ); + let plan = EgressPlan::for_request(&req).expect("a block with no port is enforceable"); + + assert!(plan.admits_udp(FORWARDER, 53)); + } + + /// A rule that names only TCP leaves DNS dropped, so matching on the + /// address alone would miss the case the warning is for. + #[test] + fn a_tcp_only_rule_does_not_admit_the_forwarder() { + let req = directional_rules_request( + NetworkAction::Deny, + vec![rule( + vec![peer("10.0.2.3/32", &[])], + vec![port(NetworkProtocol::Tcp, Some(53), None)], + )], + Vec::new(), + ); + let plan = EgressPlan::for_request(&req).expect("a tcp/53 allow is enforceable"); + + assert!(!plan.admits_udp(FORWARDER, 53)); + } + + /// Denies precede allows in the chain, so a deny covering the forwarder + /// wins even under an open default. + #[test] + fn a_deny_rule_closes_the_forwarder_under_an_open_default() { + let plan = EgressPlan::for_request(&request(NetworkAction::Allow, &[], &["10.0.2.0/24"])) + .expect("a deny rule under an open default is enforceable"); + + assert!(!plan.admits_udp(FORWARDER, 53)); + } + + /// An exclusion carves the forwarder out of a block that otherwise covers + /// it, which only address containment can see. + #[test] + fn an_exclusion_removes_the_forwarder_from_a_covering_block() { + let req = directional_rules_request( + NetworkAction::Deny, + vec![rule( + vec![peer("10.0.0.0/8", &["10.0.2.3/32"])], + vec![port(NetworkProtocol::Any, None, None)], + )], + Vec::new(), + ); + let plan = EgressPlan::for_request(&req).expect("an exclusion is enforceable"); + + assert!(!plan.admits_udp(FORWARDER, 53)); + // The rest of the block is still open, so the exclusion is what closed it. + assert!(plan.admits_udp(Ipv4Addr::new(10, 0, 2, 4), 53)); + } + + /// A port range is inclusive at both ends. + #[test] + fn a_port_range_admits_its_whole_span() { + let req = directional_rules_request( + NetworkAction::Deny, + vec![rule( + vec![peer("10.0.2.3/32", &[])], + vec![port(NetworkProtocol::Udp, Some(50), Some(60))], + )], + Vec::new(), + ); + let plan = EgressPlan::for_request(&req).expect("a port range is enforceable"); + + assert!(plan.admits_udp(FORWARDER, 50)); + assert!(plan.admits_udp(FORWARDER, 60)); + assert!(!plan.admits_udp(FORWARDER, 61)); + } + + /// The proxy posture opens one TCP endpoint, so it never admits DNS. The + /// resolver pin is not applied there, and this is what keeps the two + /// statements consistent. + #[test] + fn the_proxy_plan_does_not_admit_the_forwarder() { + let plan = EgressPlan::for_proxy(Ipv4Addr::new(10, 0, 2, 2), 3128); + + assert!(!plan.admits_udp(FORWARDER, 53)); + } + #[test] fn a_v6_destination_with_several_ports_is_reported_once() { let req = directional_rules_request( diff --git a/src/mxc-sdk/src/backends/bubblewrap/common/proxy_network.rs b/src/mxc-sdk/src/backends/bubblewrap/common/proxy_network.rs index a9b6717ef..ee12728ac 100644 --- a/src/mxc-sdk/src/backends/bubblewrap/common/proxy_network.rs +++ b/src/mxc-sdk/src/backends/bubblewrap/common/proxy_network.rs @@ -16,6 +16,7 @@ use std::thread; use std::time::{Duration, Instant}; use crate::mxc_common::filesystem_resolve::{resolve_mount_order, FsIntent}; +use crate::mxc_common::filesystem_symlink::resolve_through_symlinks; use crate::mxc_common::logger::Logger; use crate::mxc_common::models::{ContainerPolicy, ProxyAddress, ProxyHostPin}; use nix::errno::Errno; @@ -107,6 +108,19 @@ pub(crate) const SLIRP_HOST_GATEWAY_IP: Ipv4Addr = Ipv4Addr::new(10, 0, 2, 2); const SLIRP_NETWORK: &str = "10.0.2.0/24"; /// Path the hosts-file pin is mounted over inside the sandbox. const SANDBOX_HOSTS_PATH: &str = "/etc/hosts"; +/// Resolver file the sandbox reads, and the start of the resolver pin's +/// symlink walk. +const SANDBOX_RESOLV_CONF_PATH: &str = "/etc/resolv.conf"; +/// slirp's built-in DNS forwarder. +/// +/// libslirp rewrites a query addressed here to the first IPv4 nameserver in the +/// *host's* `/etc/resolv.conf` and sends it from the host's network namespace, +/// which is what puts a host loopback resolver back in reach of the sandbox. +const SLIRP_DNS_FORWARDER: &str = "10.0.2.3"; +/// The forwarder as an address, for the egress check that has to decide whether +/// the chain admits it. Kept in step with [`SLIRP_DNS_FORWARDER`] by +/// [`tests::forwarder_constants_agree`]. +pub(crate) const SLIRP_DNS_FORWARDER_IP: Ipv4Addr = Ipv4Addr::new(10, 0, 2, 3); /// Egress chain installed inside the sandbox's own network namespace. const EGRESS_CHAIN: &str = "MXC_EGRESS"; /// Chain carrying the inbound posture, hooked into `INPUT`. @@ -697,6 +711,9 @@ pub(crate) struct ProxyNetworkNamespace { userns: Option, /// Hosts file mounted over `/etc/hosts`, when the endpoint is a hostname. hosts: Option, + /// Generated resolver file and the sandbox path it is mounted over, when + /// the host's own nameservers are unreachable from a private namespace. + resolver: Option<(PathBuf, String)>, /// Read end of the descriptor the supervisor and slirp hold open, taken by /// the monitor that watches for the network provider dying mid-run. liveness_reader: Option, @@ -715,7 +732,8 @@ impl ProxyNetworkNamespace { /// supervisor signals readiness; anything it does not accept is dropped. /// `ingress` is the matching inbound posture. `pin` is the hosts-file entry /// the sandbox needs to agree with the plan, which only a hostname proxy - /// endpoint produces. + /// endpoint produces. `resolver` redirects the sandbox's nameserver to + /// slirp's forwarder, which only a loopback-resolver host produces. /// /// Callers reach this only after `BwrapRunner::validate` has already run /// [`probe_dependencies`], so the probe is not repeated here. @@ -723,6 +741,7 @@ impl ProxyNetworkNamespace { plan: &EgressPlan, ingress: &IngressPlan, pin: Option<&ProxyHostPin>, + resolver: Option<&ResolverPin>, logger: &mut Logger, script_timeout_ms: u32, ) -> Result { @@ -843,6 +862,8 @@ impl ProxyNetworkNamespace { None => None, }; + let resolver = stage_resolver(resolver, state_dir.path(), logger); + Ok(Self { state_dir, supervisor, @@ -850,6 +871,7 @@ impl ProxyNetworkNamespace { pid_writer: Some(pid_writer), userns: Some(userns), hosts, + resolver, liveness_reader: Some(liveness_reader), transactions, script_timeout_ms, @@ -884,6 +906,18 @@ impl ProxyNetworkNamespace { } } + if let Some((source, destination)) = &self.resolver { + let source = source + .to_str() + .ok_or_else(|| "Bubblewrap: resolver pin path is not valid UTF-8".to_string())?; + if insert_pin_bind(args, source, destination)? { + logger.log_line(&format!( + "Bubblewrap: the resolver pin overrides an earlier mount of {destination}; \ + the sandbox sees the pinned file" + )); + } + } + let runtime_args = [ "--userns".to_string(), userns.as_raw_fd().to_string(), @@ -1375,6 +1409,214 @@ fn strip_host_from_hosts(contents: &str, hostname: &str) -> String { out } +/// A generated resolver file and the sandbox path it is mounted over. +/// +/// The sandbox always runs in a private network namespace, so a host that +/// resolves through a loopback stub -- `systemd-resolved` at `127.0.0.53`, +/// `dnsmasq` at `127.0.0.1` -- hands it a nameserver address that belongs to +/// its own empty loopback. Binding the host's resolver directories makes the +/// file *readable*; it does not make the address *reachable*. +#[derive(Debug)] +pub(crate) struct ResolverPin { + destination: PathBuf, + contents: String, +} + +impl ResolverPin { + /// The pin this host needs, or `None` when its resolvers already work + /// inside a private namespace. + pub(crate) fn for_host(policy: &ContainerPolicy, logger: &mut Logger) -> Option { + Self::from_resolver(Path::new(SANDBOX_RESOLV_CONF_PATH), policy, logger) + } + + /// The pin a sandbox reading `resolver_path` needs. + /// + /// Returning `None` leaves the sandbox exactly as it is today, so every + /// case this cannot read confidently declines rather than guessing. + /// + /// Takes the path rather than reading the host's so the decision can be + /// exercised against a fixture. + pub(crate) fn from_resolver( + resolver_path: &Path, + policy: &ContainerPolicy, + logger: &mut Logger, + ) -> Option { + let shown = resolver_path.display(); + let host_contents = match fs::read_to_string(resolver_path) { + Ok(contents) => contents, + // No resolver file is not a loopback stub to rescue, and is normal + // on a host that configures DNS some other way. + Err(error) if error.kind() == ErrorKind::NotFound => return None, + Err(error) => { + logger.warning_line(&format!( + "WARNING: Bubblewrap: {shown} could not be read ({error}), so the sandbox \ + keeps whatever resolver it inherits and names may not resolve inside it." + )); + return None; + } + }; + if !resolvers_are_loopback_only(&host_contents) { + return None; + } + + let destination = match resolver_target_of(resolver_path) { + Ok(destination) => destination, + Err(reason) => { + logger.warning_line(&format!( + "WARNING: Bubblewrap: the host resolves through a loopback nameserver, which \ + the sandbox's private network namespace cannot reach, and {reason}. Names \ + will not resolve inside the sandbox; point {shown} at a routable nameserver." + )); + return None; + } + }; + + let destination_text = destination.to_str()?.to_string(); + if path_is_denied(policy, &destination_text) { + logger.warning_line(&format!( + "WARNING: Bubblewrap: the host resolves through a loopback nameserver, but the \ + filesystem policy denies {destination_text}, so the sandbox keeps the host's \ + unreachable resolver. Names will not resolve inside the sandbox; remove that \ + path from deniedPaths to let the sandbox use slirp's forwarder." + )); + return None; + } + + // Naming the resolver itself is the caller supplying their own, which + // the pin would otherwise mount over. + if policy_names_path(policy, &resolver_path.to_string_lossy()) + || policy_names_path(policy, &destination_text) + { + logger.log_line( + "Bubblewrap: the filesystem policy mounts its own resolver, so the sandbox keeps \ + it rather than slirp's forwarder", + ); + return None; + } + + Some(Self { + destination, + contents: render_pinned_resolv_conf(&host_contents), + }) + } + + /// Write the generated file into `directory` and report the mount it needs. + fn stage(&self, directory: &Path) -> Result<(PathBuf, String), String> { + let source = directory.join("resolv.conf"); + fs::write(&source, &self.contents) + .map_err(|error| format!("Bubblewrap: failed to write the resolver pin: {error}"))?; + let destination = self + .destination + .to_str() + .ok_or_else(|| "Bubblewrap: resolver pin path is not valid UTF-8".to_string())?; + Ok((source, destination.to_string())) + } +} + +/// Write the pin into `directory` and report the mount it needs, or `None` +/// when there is nothing to pin or staging failed. +/// +/// A staging failure gives up the pin rather than the run. The pin is an +/// improvement on a resolver the sandbox already cannot reach, so refusing to +/// start would deny the caller a sandbox that would otherwise have run with +/// exactly the DNS it has today. +fn stage_resolver( + resolver: Option<&ResolverPin>, + directory: &Path, + logger: &mut Logger, +) -> Option<(PathBuf, String)> { + let resolver = resolver?; + match resolver.stage(directory) { + Ok(staged) => { + logger.log_line(&format!( + "Bubblewrap: the host resolves through a loopback nameserver, so the sandbox \ + reads {} with nameserver {SLIRP_DNS_FORWARDER}", + staged.1 + )); + Some(staged) + } + Err(error) => { + logger.warning_line(&format!( + "WARNING: {error}. The sandbox keeps the host's loopback nameserver, which its \ + private network namespace cannot reach, so names will not resolve inside it." + )); + None + } + } +} + +/// Whether every nameserver the host declares is a loopback address, with at +/// least one of them IPv4. +/// +/// Both halves gate the pin. All-loopback means name resolution is already +/// broken inside the namespace, so replacing it cannot regress a host that +/// works. The IPv4 requirement is libslirp's: it forwards a query to the first +/// IPv4 nameserver in the host's file and drops the packet when there is none. +fn resolvers_are_loopback_only(contents: &str) -> bool { + let mut saw_ipv4_loopback = false; + for value in nameserver_values(contents) { + // A link-local v6 nameserver may carry a zone suffix, which is not part + // of the address. + let literal = value.split('%').next().unwrap_or(value); + match literal.parse::() { + Ok(address) if address.is_loopback() => saw_ipv4_loopback |= address.is_ipv4(), + _ => return false, + } + } + saw_ipv4_loopback +} + +/// The sandbox's resolver file, pointed at slirp's forwarder. +/// +/// Every other directive is carried over: `search` and `options` decide how a +/// bare name is expanded and retried, and dropping them would change which +/// names resolve rather than merely where they resolve. +fn render_pinned_resolv_conf(contents: &str) -> String { + let mut out = format!("nameserver {SLIRP_DNS_FORWARDER}\n"); + for line in contents.lines().filter(|line| !is_nameserver_line(line)) { + out.push_str(line); + out.push('\n'); + } + out +} + +fn nameserver_values(contents: &str) -> impl Iterator { + contents + .lines() + .filter(|line| is_nameserver_line(line)) + .filter_map(|line| line.split_whitespace().nth(1)) +} + +fn is_nameserver_line(line: &str) -> bool { + // A resolver directive is only read when its keyword starts the line. + !line.starts_with(char::is_whitespace) && line.split_whitespace().next() == Some("nameserver") +} + +/// Where the pin has to be mounted for the sandbox to read it. +/// +/// bwrap refuses to mount over a path whose leaf is a symlink, and cannot +/// create a mount point beneath a symlinked ancestor that sits inside a +/// read-only bind, so the destination is resolved the same way a `deniedPaths` +/// entry is: every component, not only the leaf. Resolving less than that also +/// made the policy checks compare a spelling the mount would never use. +fn resolver_target_of(start: &Path) -> Result { + let resolved = resolve_through_symlinks(start) + .ok_or_else(|| format!("{} is not a resolvable path", start.display()))?; + + // A chain that loops or dangles leaves canonicalization unfinished, so the + // walk hands back path text that is still a link. Binding there aborts the + // sandbox, which is worse than the broken DNS being fixed. + let metadata = fs::symlink_metadata(&resolved) + .map_err(|error| format!("{} could not be read ({error})", resolved.display()))?; + if !metadata.file_type().is_file() { + return Err(format!( + "{} does not resolve to a regular file", + resolved.display() + )); + } + Ok(resolved) +} + /// Reject a hostname endpoint whose pin would defeat a denied `/etc/hosts`. /// /// The pin is spliced after every filesystem-policy mount so it survives them @@ -1412,28 +1654,33 @@ pub(crate) fn check_hosts_pin_against_policy( /// path, so the deepest entry covering the file is the one that takes effect: /// an ancestor denial counts, and a more specific grant beneath it wins back. fn hosts_file_is_denied(policy: &ContainerPolicy) -> bool { + path_is_denied(policy, SANDBOX_HOSTS_PATH) +} + +/// Whether the filesystem policy masks `target`. +fn path_is_denied(policy: &ContainerPolicy, target: &str) -> bool { resolve_mount_order(policy) .iter() - .rfind(|mount| covers_hosts_file(&mount.path)) + .rfind(|mount| covers_path(&mount.path, target)) .is_some_and(|mount| mount.intent == FsIntent::Denied) } -/// Whether `path` is the sandbox hosts file or a directory holding it. -fn covers_hosts_file(path: &str) -> bool { +/// Whether `path` is `target` or a directory holding it. +fn covers_path(path: &str, target: &str) -> bool { let path = path.trim_end_matches('/'); - path == SANDBOX_HOSTS_PATH - || SANDBOX_HOSTS_PATH + path == target + || target .strip_prefix(path) .is_some_and(|rest| rest.starts_with('/')) } -/// Splice the pinned-hosts bind in just before the command separator. -/// -/// bwrap applies mounts in argument order and the last mount at a path wins, -/// so the bind must come after every baseline and user-policy mount for the -/// pin to survive -- including one that would otherwise expose the host's own -/// `/etc/hosts`. Returns `true` when an earlier mount already targeted -/// `/etc/hosts`, so the caller can report that the pin overrides it. +/// Whether the filesystem policy names `target` itself rather than an ancestor. +fn policy_names_path(policy: &ContainerPolicy, target: &str) -> bool { + resolve_mount_order(policy) + .iter() + .any(|mount| mount.path.trim_end_matches('/') == target) +} + /// Index of the separator that ends bwrap's options and begins the command. /// /// Scanning for the first `--` would find a caller-controlled value instead: @@ -1449,22 +1696,35 @@ fn command_separator(args: &[String]) -> Result { .ok_or_else(|| "Bubblewrap: argument list has no command separator".to_string()) } -fn insert_hosts_bind(args: &mut Vec, hosts_path: &str) -> Result { +/// Splice a pin's bind in just before the command separator. +/// +/// bwrap applies mounts in argument order and the last mount at a path wins, +/// so the bind must come after every baseline and user-policy mount for the +/// pin to survive -- including one that would otherwise expose the host's own +/// file. Returns `true` when an earlier mount already targeted `destination`, +/// so the caller can report that the pin overrides it. +fn insert_pin_bind( + args: &mut Vec, + source: &str, + destination: &str, +) -> Result { let separator = command_separator(args)?; - let overrides = args[..separator] - .iter() - .any(|arg| arg == SANDBOX_HOSTS_PATH); + let overrides = args[..separator].iter().any(|arg| arg == destination); args.splice( separator..separator, [ "--ro-bind".to_string(), - hosts_path.to_string(), - SANDBOX_HOSTS_PATH.to_string(), + source.to_string(), + destination.to_string(), ], ); Ok(overrides) } +fn insert_hosts_bind(args: &mut Vec, hosts_path: &str) -> Result { + insert_pin_bind(args, hosts_path, SANDBOX_HOSTS_PATH) +} + /// What pulled this request into a private network namespace. /// /// The dependency probe is shared by runtime proxy egress and directional @@ -2590,6 +2850,520 @@ mod tests { assert!(insert_hosts_bind(&mut args, "/tmp/pin/hosts").is_err()); } + /// The stub resolver the issue reports: readable inside the sandbox, + /// unreachable from it. + const SYSTEMD_RESOLVED_STUB: &str = "nameserver 127.0.0.53\noptions edns0 trust-ad\n\ + search corp.example\n"; + + /// The resolver file is written from the string and the egress check reads + /// the address. A drift between them would pin one forwarder and test the + /// chain against another, so the warning would fire on a chain that works. + #[test] + fn forwarder_constants_agree() { + assert_eq!(SLIRP_DNS_FORWARDER_IP.to_string(), SLIRP_DNS_FORWARDER); + } + + #[test] + fn a_loopback_stub_resolver_is_pinned() { + assert!(resolvers_are_loopback_only(SYSTEMD_RESOLVED_STUB)); + assert!(resolvers_are_loopback_only("nameserver 127.0.0.1\n")); + } + + /// Pinning a host whose nameserver is routable would redirect name + /// resolution that already works. + #[test] + fn a_routable_resolver_is_left_alone() { + assert!(!resolvers_are_loopback_only("nameserver 8.8.8.8\n")); + assert!(!resolvers_are_loopback_only( + "nameserver 127.0.0.53\nnameserver 8.8.8.8\n" + )); + } + + /// slirp forwards to the first IPv4 nameserver in the host's file and drops + /// the query when there is none, so a v6-only resolver list cannot be + /// rescued by pointing at the forwarder. + #[test] + fn a_resolver_list_without_an_ipv4_entry_is_left_alone() { + assert!(!resolvers_are_loopback_only("nameserver ::1\n")); + assert!(!resolvers_are_loopback_only("search corp.example\n")); + assert!(!resolvers_are_loopback_only("")); + } + + /// An address this cannot parse might be routable, and pinning over it + /// would take away name resolution that works. + #[test] + fn an_unreadable_nameserver_is_left_alone() { + assert!(!resolvers_are_loopback_only("nameserver localhost\n")); + assert!(!resolvers_are_loopback_only("nameserver 127.0.0.999\n")); + assert!(!resolvers_are_loopback_only("nameserver\n")); + } + + /// glibc reads a directive only when its keyword starts the line, so an + /// indented one names no resolver the host actually uses and the file still + /// counts as loopback-only. + #[test] + fn an_indented_directive_is_not_a_nameserver() { + assert!(!is_nameserver_line(" nameserver 8.8.8.8")); + assert!(resolvers_are_loopback_only( + "nameserver 127.0.0.53\n nameserver 8.8.8.8\n" + )); + assert_eq!( + render_pinned_resolv_conf("nameserver 127.0.0.53\n nameserver 8.8.8.8\n"), + "nameserver 10.0.2.3\n nameserver 8.8.8.8\n" + ); + } + + /// glibc takes the address after the keyword and ignores the rest of the + /// line. + #[test] + fn a_nameserver_line_ignores_trailing_fields() { + assert!(resolvers_are_loopback_only( + "nameserver 127.0.0.53 # stub\n" + )); + } + + /// `search` and `options` decide how a bare name is expanded and retried, + /// so dropping them would change which names resolve. + #[test] + fn the_pinned_resolver_keeps_every_other_directive() { + let pinned = render_pinned_resolv_conf(SYSTEMD_RESOLVED_STUB); + + assert_eq!( + pinned, + "nameserver 10.0.2.3\noptions edns0 trust-ad\nsearch corp.example\n" + ); + } + + /// Leaving one behind would let the resolver fall back to an address the + /// namespace cannot reach. + #[test] + fn the_pinned_resolver_replaces_every_nameserver() { + let pinned = + render_pinned_resolv_conf("nameserver 127.0.0.53\nnameserver 127.0.0.54\n# comment\n"); + + assert_eq!(pinned, "nameserver 10.0.2.3\n# comment\n"); + assert_eq!(nameserver_values(&pinned).collect::>(), ["10.0.2.3"]); + } + + /// A host whose `/etc/resolv.conf` is `contents`, as a real file the + /// decision can be run against. + fn host_with_resolver(dir: &tempfile::TempDir, contents: &str) -> PathBuf { + let root = std::fs::canonicalize(dir.path()).expect("tempdir canonicalizes"); + let path = root.join("resolv.conf"); + fs::write(&path, contents).expect("fixture resolver"); + path + } + + fn buffer_logger() -> Logger { + Logger::new(crate::mxc_common::logger::Mode::Buffer) + } + + /// The whole decision, not just its helpers: a loopback stub is pinned to + /// slirp's forwarder and mounted over the file the chain ends at. + #[test] + fn the_decision_pins_a_loopback_host() { + let dir = tempfile::tempdir().unwrap(); + let resolver = host_with_resolver(&dir, SYSTEMD_RESOLVED_STUB); + let mut logger = buffer_logger(); + + let pin = ResolverPin::from_resolver(&resolver, &ContainerPolicy::default(), &mut logger) + .expect("a loopback stub is pinned"); + + assert_eq!(pin.destination, resolver); + assert!(pin.contents.contains("nameserver 10.0.2.3")); + assert!(pin.contents.contains("search corp.example")); + assert!(logger.warnings().is_empty(), "{:?}", logger.warnings()); + } + + /// A host that already resolves is left alone, silently. + #[test] + fn the_decision_declines_a_routable_host() { + let dir = tempfile::tempdir().unwrap(); + let resolver = host_with_resolver(&dir, "nameserver 8.8.8.8\n"); + let mut logger = buffer_logger(); + + assert!( + ResolverPin::from_resolver(&resolver, &ContainerPolicy::default(), &mut logger) + .is_none() + ); + assert!(logger.warnings().is_empty(), "{:?}", logger.warnings()); + } + + /// A host with no resolver file has no stub to rescue, which is not a + /// degradation worth reporting. + #[test] + fn the_decision_is_silent_when_there_is_no_resolver() { + let dir = tempfile::tempdir().unwrap(); + let absent = dir.path().join("absent.conf"); + let mut logger = buffer_logger(); + + assert!( + ResolverPin::from_resolver(&absent, &ContainerPolicy::default(), &mut logger).is_none() + ); + assert!(logger.warnings().is_empty(), "{:?}", logger.warnings()); + } + + /// The pin is spliced after every policy mount, so honouring it would hand + /// back the file the policy masked. + #[test] + fn the_decision_declines_a_denied_destination() { + let dir = tempfile::tempdir().unwrap(); + let resolver = host_with_resolver(&dir, SYSTEMD_RESOLVED_STUB); + let policy = ContainerPolicy { + denied_paths: vec![resolver.to_string_lossy().into_owned()], + ..Default::default() + }; + let mut logger = buffer_logger(); + + assert!(ResolverPin::from_resolver(&resolver, &policy, &mut logger).is_none()); + let out = logger.warnings().join("\n"); + assert!(out.contains("deniedPaths"), "must name the remedy: {out}"); + } + + #[test] + fn the_decision_declines_a_caller_supplied_resolver() { + let dir = tempfile::tempdir().unwrap(); + let resolver = host_with_resolver(&dir, SYSTEMD_RESOLVED_STUB); + let policy = ContainerPolicy { + readonly_paths: vec![resolver.to_string_lossy().into_owned()], + ..Default::default() + }; + let mut logger = buffer_logger(); + + assert!(ResolverPin::from_resolver(&resolver, &policy, &mut logger).is_none()); + assert!( + logger.warnings().is_empty(), + "the caller got what they asked for: {:?}", + logger.warnings() + ); + } + + /// A chain with no regular file at the end gives up the pin, not the run, + /// and says why. + #[test] + fn the_decision_declines_an_unresolvable_chain() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + let first = root.join("first.conf"); + let second = root.join("second.conf"); + std::os::unix::fs::symlink(&second, &first).unwrap(); + std::os::unix::fs::symlink(&first, &second).unwrap(); + let mut logger = buffer_logger(); + + assert!( + ResolverPin::from_resolver(&first, &ContainerPolicy::default(), &mut logger).is_none() + ); + let out = logger.warnings().join("\n"); + assert!(out.contains("could not be read"), "got: {out}"); + } + + /// The destination is the canonical file, so a policy denying it through an + /// ancestor is honoured even when the host's symlink spells the path + /// differently. An ancestor denial is used deliberately: a policy naming + /// the destination exactly would also trip the caller-supplied-resolver + /// check, and the test would pass without the deny comparison being right. + #[test] + fn the_decision_compares_policy_against_the_canonical_destination() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + fs::create_dir_all(root.join("real")).unwrap(); + fs::write(root.join("real/stub.conf"), SYSTEMD_RESOLVED_STUB).unwrap(); + std::os::unix::fs::symlink(root.join("real"), root.join("link")).unwrap(); + + let policy = ContainerPolicy { + denied_paths: vec![root.join("real").to_string_lossy().into_owned()], + ..Default::default() + }; + let mut logger = buffer_logger(); + + assert!( + ResolverPin::from_resolver(&root.join("link/stub.conf"), &policy, &mut logger) + .is_none(), + "a denial of the canonical directory must stop the pin reached by an alias" + ); + let out = logger.warnings().join("\n"); + assert!(out.contains("deniedPaths"), "must name the remedy: {out}"); + } + + /// The pin improves on a resolver the sandbox already cannot reach, so a + /// staging failure must cost the pin and not the run — the caller would + /// otherwise lose a sandbox that was going to start. + #[test] + fn a_staging_failure_gives_up_the_pin_rather_than_the_run() { + let dir = tempfile::tempdir().unwrap(); + let resolver = host_with_resolver(&dir, SYSTEMD_RESOLVED_STUB); + let mut logger = buffer_logger(); + let pin = ResolverPin::from_resolver(&resolver, &ContainerPolicy::default(), &mut logger) + .expect("a loopback stub is pinned"); + + let staged = stage_resolver( + Some(&pin), + Path::new("/mxc-absent-staging-dir"), + &mut logger, + ); + + assert!( + staged.is_none(), + "a failed stage leaves the sandbox unpinned" + ); + let out = logger.warnings().join("\n"); + assert!( + out.contains("loopback nameserver"), + "the degradation must be reported: {out}" + ); + assert!( + logger.get_buffer().is_empty(), + "the warning must travel as a retained warning, not as buffer output" + ); + } + + #[test] + fn staging_reports_the_pin_it_mounted() { + let dir = tempfile::tempdir().unwrap(); + let resolver = host_with_resolver(&dir, SYSTEMD_RESOLVED_STUB); + let mut logger = buffer_logger(); + let pin = ResolverPin::from_resolver(&resolver, &ContainerPolicy::default(), &mut logger) + .expect("a loopback stub is pinned"); + + let staging = tempfile::tempdir().unwrap(); + let (source, destination) = stage_resolver(Some(&pin), staging.path(), &mut logger) + .expect("staging into a real dir"); + + assert_eq!(fs::read_to_string(&source).unwrap(), pin.contents); + assert_eq!(destination, resolver.to_string_lossy()); + assert!( + logger.warnings().is_empty(), + "a successful pin is not a degradation: {:?}", + logger.warnings() + ); + } + + /// A host whose resolver already works is never pinned, and that is not a + /// degradation worth reporting. + #[test] + fn staging_nothing_is_silent() { + let dir = tempfile::tempdir().unwrap(); + let mut logger = buffer_logger(); + + assert!(stage_resolver(None, dir.path(), &mut logger).is_none()); + assert!(logger.warnings().is_empty()); + assert!(logger.get_buffer().is_empty()); + } + + /// bwrap refuses to mount over a symlinked leaf, so the pin has to land on + /// the file the chain ends at. + #[test] + fn the_resolver_target_follows_the_link_chain() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + fs::create_dir_all(root.join("run/systemd/resolve")).unwrap(); + fs::create_dir_all(root.join("etc")).unwrap(); + let stub = root.join("run/systemd/resolve/stub-resolv.conf"); + fs::write(&stub, SYSTEMD_RESOLVED_STUB).unwrap(); + std::os::unix::fs::symlink(&stub, root.join("etc/resolv.conf")).unwrap(); + + assert_eq!(resolver_target_of(&root.join("etc/resolv.conf")), Ok(stub)); + } + + /// Debian's resolvconf writes a relative link, which has to be resolved + /// against the directory holding it rather than the process's cwd. + #[test] + fn the_resolver_target_resolves_a_relative_link() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + fs::create_dir_all(root.join("run/resolvconf")).unwrap(); + fs::create_dir_all(root.join("etc")).unwrap(); + let real = root.join("run/resolvconf/resolv.conf"); + fs::write(&real, SYSTEMD_RESOLVED_STUB).unwrap(); + std::os::unix::fs::symlink( + "../run/resolvconf/resolv.conf", + root.join("etc/resolv.conf"), + ) + .unwrap(); + + assert_eq!(resolver_target_of(&root.join("etc/resolv.conf")), Ok(real)); + } + + /// A destination still holding a symlinked ancestor is one bwrap cannot + /// create a mount point at when that ancestor's parent is bound read-only, + /// which aborts the sandbox rather than merely leaving DNS broken. + #[test] + fn the_resolver_target_resolves_a_symlinked_ancestor() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + fs::create_dir_all(root.join("run/res/v1")).unwrap(); + fs::create_dir_all(root.join("etc")).unwrap(); + let real = root.join("run/res/v1/stub.conf"); + fs::write(&real, SYSTEMD_RESOLVED_STUB).unwrap(); + std::os::unix::fs::symlink(root.join("run/res/v1"), root.join("run/res/current")).unwrap(); + std::os::unix::fs::symlink( + root.join("run/res/current/stub.conf"), + root.join("etc/resolv.conf"), + ) + .unwrap(); + + assert_eq!(resolver_target_of(&root.join("etc/resolv.conf")), Ok(real)); + } + + /// `..` has to be applied to the directory the link really sits in, not to + /// path text that still names a symlink, or the pin covers a file the + /// resolver never reads. + #[test] + fn the_resolver_target_folds_dotdot_under_a_symlinked_ancestor() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + fs::create_dir_all(root.join("run/resolvconf")).unwrap(); + fs::create_dir_all(root.join("etc")).unwrap(); + let real = root.join("run/other.conf"); + fs::write(&real, SYSTEMD_RESOLVED_STUB).unwrap(); + // /etc/link -> /run/resolvconf, then the leaf climbs out of it. + std::os::unix::fs::symlink(root.join("run/resolvconf"), root.join("etc/link")).unwrap(); + std::os::unix::fs::symlink( + root.join("etc/link/../other.conf"), + root.join("etc/resolv.conf"), + ) + .unwrap(); + + assert_eq!(resolver_target_of(&root.join("etc/resolv.conf")), Ok(real)); + } + + /// A plain file is its own target, and a missing one has nothing to pin + /// over. + #[test] + fn the_resolver_target_handles_a_regular_file_and_a_broken_link() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + let regular = root.join("resolv.conf"); + fs::write(®ular, SYSTEMD_RESOLVED_STUB).unwrap(); + assert_eq!(resolver_target_of(®ular), Ok(regular.clone())); + + let dangling = root.join("dangling"); + std::os::unix::fs::symlink(root.join("absent"), &dangling).unwrap(); + assert!(resolver_target_of(&dangling).is_err()); + } + + /// A link loop leaves canonicalization unfinished, so the walk hands back a + /// path that is still a link; binding there would abort the sandbox. + #[test] + fn the_resolver_target_gives_up_on_a_link_loop() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + let first = root.join("first"); + let second = root.join("second"); + std::os::unix::fs::symlink(&second, &first).unwrap(); + std::os::unix::fs::symlink(&first, &second).unwrap(); + + let error = resolver_target_of(&first).expect_err("a loop has no target"); + assert!(error.contains("regular file"), "got: {error}"); + } + + /// A directory would be mounted over as if it were the resolver file. + #[test] + fn the_resolver_target_refuses_a_directory() { + let dir = tempfile::tempdir().unwrap(); + let root = std::fs::canonicalize(dir.path()).unwrap(); + fs::create_dir(root.join("adir")).unwrap(); + + assert!(resolver_target_of(&root.join("adir")).is_err()); + } + + /// The pin is spliced after every policy mount, so a `deniedPaths` entry + /// covering the resolver would be handed back a readable file instead. + #[test] + fn a_denied_resolver_path_is_detected() { + let denied = ContainerPolicy { + denied_paths: vec!["/run/systemd/resolve".into()], + ..Default::default() + }; + assert!(path_is_denied( + &denied, + "/run/systemd/resolve/stub-resolv.conf" + )); + assert!(!path_is_denied(&denied, "/etc/resolv.conf")); + + let regranted = ContainerPolicy { + denied_paths: vec!["/run/systemd/resolve".into()], + readonly_paths: vec!["/run/systemd/resolve/stub-resolv.conf".into()], + ..Default::default() + }; + assert!(!path_is_denied( + ®ranted, + "/run/systemd/resolve/stub-resolv.conf" + )); + } + + /// A caller who mounts their own resolver has chosen one, and the pin is + /// spliced after every policy mount, so it would replace their choice. + #[test] + fn a_caller_supplied_resolver_is_recognized() { + let supplied = ContainerPolicy { + readonly_paths: vec!["/etc/resolv.conf".into()], + ..Default::default() + }; + assert!(policy_names_path(&supplied, "/etc/resolv.conf")); + + // An ancestor grant is not a choice about the resolver itself. + let whole_etc = ContainerPolicy { + readonly_paths: vec!["/etc".into()], + ..Default::default() + }; + assert!(!policy_names_path(&whole_etc, "/etc/resolv.conf")); + } + + /// The pin only works if it outranks the baseline mount of the directory + /// holding the resolver. + #[test] + fn the_resolver_bind_lands_after_the_policy_mounts() { + let mut args = vec![ + "--ro-bind-try".to_string(), + "/run/systemd/resolve".to_string(), + "/run/systemd/resolve".to_string(), + ]; + args.extend(command_tail()); + + let overrides = insert_pin_bind( + &mut args, + "/tmp/pin/resolv.conf", + "/run/systemd/resolve/stub-resolv.conf", + ) + .unwrap(); + + assert!(!overrides, "a parent directory mount is not the same path"); + let pin = args + .windows(3) + .position(|window| { + window + == [ + "--ro-bind", + "/tmp/pin/resolv.conf", + "/run/systemd/resolve/stub-resolv.conf", + ] + }) + .expect("pin bind should be present"); + assert!(pin > 0, "pin must follow the baseline mount"); + assert_eq!( + args[args.len() - COMMAND_TAIL.len() - 1..], + ["--", "sh", "-c", "echo hello"] + ); + } + + /// Reported so an operator who mounted their own resolver knows the pin + /// shadowed it. + #[test] + fn the_resolver_bind_reports_overriding_an_earlier_mount() { + let mut args = vec![ + "--ro-bind".to_string(), + "/custom/resolv.conf".to_string(), + "/etc/resolv.conf".to_string(), + ]; + args.extend(command_tail()); + + assert!( + insert_pin_bind(&mut args, "/tmp/pin/resolv.conf", "/etc/resolv.conf").unwrap(), + "an earlier mount of the same path is an override" + ); + } + /// The command bwrap is asked to run, as `build_args` appends it. fn command_tail() -> Vec { COMMAND_TAIL @@ -4938,6 +5712,7 @@ exec sleep 30 pid_writer: None, userns: None, hosts: None, + resolver: None, liveness_reader: None, transactions: 0, script_timeout_ms: 0, diff --git a/src/mxc-sdk/src/core/mxc_common/filesystem_symlink.rs b/src/mxc-sdk/src/core/mxc_common/filesystem_symlink.rs new file mode 100644 index 000000000..65773d218 --- /dev/null +++ b/src/mxc-sdk/src/core/mxc_common/filesystem_symlink.rs @@ -0,0 +1,162 @@ +// Copyright (c) Microsoft Corporation. +// Licensed under the MIT License. + +//! Symlink resolution for host paths the Linux backends mount over. +//! +//! The Linux containment backends realise filesystem policy as bind mounts, and +//! bwrap cannot create a mount point when any component of the destination — +//! the leaf or an ancestor directory — is a pre-existing host symlink whose +//! parent is bound into the sandbox: the mount resolves through the host +//! symlink and fails with `ENOENT`, aborting the sandbox. Every path a backend +//! intends to mount over is therefore resolved to its real location first. +//! +//! Like [`crate::mxc_common::filesystem_object`] this does file I/O, so it +//! lives in `mxc_common` and is invoked by backend runners close to the point +//! of enforcement. + +use std::path::{Component, Path, PathBuf}; + +/// Resolve every symlink in `path` (leaf and ancestors) to a real filesystem +/// path, tolerating trailing components that do not exist yet. +/// +/// `std::fs::canonicalize` resolves symlinks at every level but requires the +/// **whole** path to exist. To also cover not-yet-created denied paths under a +/// symlinked ancestor, this walks the components from the root: every existing +/// prefix is canonicalized (following symlinks exactly like the kernel), while +/// `.` and `..` in the not-yet-existent tail are folded lexically. Folding `..` +/// this way is safe because a component that does not exist cannot be a symlink, +/// so the result matches the target the kernel's path resolution would reach. +/// Returns `None` only for an empty path. +/// +/// A chain that loops or dangles leaves canonicalization unable to finish, so +/// the result is the lexical text and is still a symlink. Callers that are +/// about to mount over the result must check what they got. +/// +/// A naive backward walk that collected `file_name()` silently dropped `..` +/// components (Rust returns `None` for a `..` file name) and reconstructed the +/// wrong target: `/link/missing/../secret` became `/real/missing/secret` +/// instead of `/real/secret`, so the mask landed on a bystander path and the +/// real denied target stayed exposed. +pub(crate) fn resolve_through_symlinks(path: &Path) -> Option { + let mut result = PathBuf::new(); + for component in path.components() { + match component { + Component::Prefix(_) | Component::RootDir => result.push(component.as_os_str()), + Component::CurDir => {} + Component::ParentDir => { + result.pop(); + } + Component::Normal(name) => { + result.push(name); + // Canonicalize the prefix so far so symlinks are followed while + // it still exists; once a component is missing, canonicalize + // fails and the remaining tail is folded lexically above. + if let Ok(real) = std::fs::canonicalize(&result) { + result = real; + } + } + } + } + if result.as_os_str().is_empty() { + None + } else { + Some(result) + } +} + +#[cfg(test)] +mod tests { + use super::*; + + /// Canonicalized up front so a symlinked `TMPDIR` does not make every + /// expectation below differ from what the walk returns. + fn temp_root(dir: &tempfile::TempDir) -> PathBuf { + std::fs::canonicalize(dir.path()).expect("tempdir canonicalizes") + } + + #[test] + fn a_leaf_symlink_resolves_to_its_target() { + let dir = tempfile::tempdir().unwrap(); + let root = temp_root(&dir); + let target = root.join("real.conf"); + std::fs::write(&target, b"x").unwrap(); + let link = root.join("link.conf"); + std::os::unix::fs::symlink(&target, &link).unwrap(); + + assert_eq!(resolve_through_symlinks(&link), Some(target)); + } + + /// The case a leaf-only walk misses: the destination still names a symlink + /// directory, which is where bwrap refuses to create the mount point. + #[test] + fn a_symlinked_ancestor_resolves() { + let dir = tempfile::tempdir().unwrap(); + let root = temp_root(&dir); + std::fs::create_dir_all(root.join("real/sub")).unwrap(); + let file = root.join("real/sub/f.conf"); + std::fs::write(&file, b"x").unwrap(); + std::os::unix::fs::symlink(root.join("real"), root.join("link")).unwrap(); + + assert_eq!( + resolve_through_symlinks(&root.join("link/sub/f.conf")), + Some(file) + ); + } + + /// `..` must be applied to the directory the symlink really points at. + #[test] + fn dotdot_folds_against_the_resolved_parent() { + let dir = tempfile::tempdir().unwrap(); + let root = temp_root(&dir); + std::fs::create_dir(root.join("real")).unwrap(); + let sibling = root.join("sibling.conf"); + std::fs::write(&sibling, b"x").unwrap(); + std::os::unix::fs::symlink(root.join("real"), root.join("link")).unwrap(); + + assert_eq!( + resolve_through_symlinks(&root.join("link/../sibling.conf")), + Some(sibling) + ); + } + + /// A path whose tail does not exist yet still resolves its real ancestors, + /// which is what lets a denied path be masked before it is created. + #[test] + fn a_missing_tail_keeps_the_resolved_ancestors() { + let dir = tempfile::tempdir().unwrap(); + let root = temp_root(&dir); + std::fs::create_dir(root.join("real")).unwrap(); + std::os::unix::fs::symlink(root.join("real"), root.join("link")).unwrap(); + + assert_eq!( + resolve_through_symlinks(&root.join("link/absent.conf")), + Some(root.join("real/absent.conf")) + ); + } + + /// Canonicalization cannot finish, so the caller is handed path text that + /// is still a link rather than an error. + #[test] + fn a_loop_yields_unresolved_text() { + let dir = tempfile::tempdir().unwrap(); + let root = temp_root(&dir); + let first = root.join("first"); + let second = root.join("second"); + std::os::unix::fs::symlink(&second, &first).unwrap(); + std::os::unix::fs::symlink(&first, &second).unwrap(); + + let resolved = resolve_through_symlinks(&first).expect("a non-empty path resolves"); + assert!( + std::fs::symlink_metadata(&resolved) + .expect("the link node exists") + .file_type() + .is_symlink(), + "a loop leaves the result a symlink: {resolved:?}" + ); + } + + #[test] + fn an_empty_path_has_no_target() { + assert_eq!(resolve_through_symlinks(Path::new("")), None); + } +} diff --git a/src/mxc-sdk/src/core/mxc_common/mod.rs b/src/mxc-sdk/src/core/mxc_common/mod.rs index abe02c2c2..0fd5d5908 100644 --- a/src/mxc-sdk/src/core/mxc_common/mod.rs +++ b/src/mxc-sdk/src/core/mxc_common/mod.rs @@ -72,6 +72,11 @@ pub mod system_dir; #[cfg(unix)] pub mod interruptible_reader; +// Linux-only: used by the backends that realise filesystem policy as bind +// mounts, which must resolve a destination before mounting over it. +#[cfg(target_os = "linux")] +pub(crate) mod filesystem_symlink; + /// Crate-wide lock and guards for tests that mutate process environment. #[cfg(all(test, target_os = "windows"))] pub(crate) mod test_env; diff --git a/tests/configs/bubblewrap_network_dns_stub.json b/tests/configs/bubblewrap_network_dns_stub.json new file mode 100644 index 000000000..87f06db87 --- /dev/null +++ b/tests/configs/bubblewrap_network_dns_stub.json @@ -0,0 +1,21 @@ +{ + "$schema": "../../schemas/stable/mxc-config.schema.0.9.0-alpha.json", + "version": "0.9.0-alpha", + "containerId": "CLI-Bubblewrap-Network-DNS-Stub", + "containment": "bubblewrap", + "process": { + "commandLine": "sh -c \"echo SANDBOX_RESOLV_BEGIN; cat /etc/resolv.conf; echo SANDBOX_RESOLV_END; getent hosts mxc-dns-probe.invalid || echo MXC_DNS_LOOKUP_FAILED\"", + "env": [ + "PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin" + ] + }, + "network": { + "egress": { + "default": "allow" + }, + "ingress": { + "default": "deny", + "hostLoopback": "deny" + } + } +} diff --git a/tests/scripts/run_bwrap_all_tests.sh b/tests/scripts/run_bwrap_all_tests.sh index 7c1a92385..3fc2d98c4 100644 --- a/tests/scripts/run_bwrap_all_tests.sh +++ b/tests/scripts/run_bwrap_all_tests.sh @@ -91,6 +91,7 @@ run_test "Bubblewrap Denied Masking" "$SCRIPT_DIR/run_bwrap_denied_masking_test. run_test "Bubblewrap Network Block" "$SCRIPT_DIR/run_bwrap_network_test.sh" run_test "Bubblewrap Network Proxy" "$SCRIPT_DIR/run_bwrap_network_proxy_test.sh" run_test "Bubblewrap Network Firewall" "$SCRIPT_DIR/run_bwrap_firewall_test.sh" +run_test "Bubblewrap Loopback Resolver DNS" "$SCRIPT_DIR/run_bwrap_dns_test.sh" run_test "Bubblewrap Directional Network" "$SCRIPT_DIR/run_bwrap_directional_test.sh" run_test "Bubblewrap allowLocalNetwork" "$SCRIPT_DIR/run_bwrap_localnet_test.sh" run_test "Bubblewrap Inbound Deny" "$SCRIPT_DIR/run_bwrap_inbound_deny_test.sh" diff --git a/tests/scripts/run_bwrap_dns_test.sh b/tests/scripts/run_bwrap_dns_test.sh new file mode 100644 index 000000000..dee2c921e --- /dev/null +++ b/tests/scripts/run_bwrap_dns_test.sh @@ -0,0 +1,202 @@ +#!/bin/bash +# Bubblewrap name-resolution test for a host that resolves through a loopback +# stub, which is the default on Ubuntu and Fedora. +# +# The bug this suite is designed to catch: the sandbox inherited the host's +# `/etc/resolv.conf` naming a loopback nameserver, which inside a private +# network namespace is the sandbox's own empty loopback, so every lookup +# failed while every numeric-address network test kept passing. +# +# The host running this does not have to resolve through a loopback stub -- +# the test builds one. It re-executes itself in a private user, mount, and +# network namespace, serves DNS on 127.0.0.53 there, and points the resolver +# at it, so the condition is identical on a developer box and on CI. +set -euo pipefail + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +REPO_DIR="$(dirname "$(dirname "$SCRIPT_DIR")")" +if [ -n "${LXC_EXEC:-}" ]; then + if [ ! -f "$LXC_EXEC" ]; then + echo "Error: LXC_EXEC is set to '$LXC_EXEC', which does not exist." + exit 1 + fi +else + LXC_EXEC="$REPO_DIR/src/target/release/lxc-exec" + if [ ! -f "$LXC_EXEC" ]; then + LXC_EXEC="$REPO_DIR/src/target/debug/lxc-exec" + fi + if [ ! -f "$LXC_EXEC" ]; then + echo "Error: lxc-exec not found. Run build.sh first." + exit 1 + fi +fi + +CONFIG="$REPO_DIR/tests/configs/bubblewrap_network_dns_stub.json" +# A reserved TLD can never resolve through a real resolver, so an answer can +# only have come from the peer below. Keep these in step with the fixture. +PROBE_NAME="mxc-dns-probe.invalid" +PROBE_ANSWER="203.0.113.99" +STUB_ADDRESS="127.0.0.53" +SEARCH_DOMAIN="mxc-test.invalid" +SLIRP_FORWARDER="10.0.2.3" + +if ! grep -Fq "$PROBE_NAME" "$CONFIG"; then + echo "FAIL: ${CONFIG##*/} no longer probes $PROBE_NAME; script and fixture drifted." + exit 1 +fi + +for tool in bwrap unshare python3 timeout ip getent; do + if ! command -v "$tool" >/dev/null 2>&1; then + echo "FAIL: $tool is required to serve and query the stub resolver." + exit 1 + fi +done +if ! command -v slirp4netns >/dev/null 2>&1; then + echo "SKIP: slirp4netns not installed; the resolver pin needs the private namespace." + exit 77 +fi + +# Everything below has to run against a loopback resolver, which this builds +# rather than requires. The namespaces are unprivileged, so the sandbox still +# launches the way it does for a real caller. +if [ "${MXC_BWRAP_DNS_INNER:-0}" != "1" ]; then + if ! unshare -Urmn true >/dev/null 2>&1; then + echo "SKIP: unprivileged user, mount, and network namespaces are unavailable;" + echo " a loopback-resolver host cannot be simulated here." + exit 77 + fi + exec unshare -Urmn env MXC_BWRAP_DNS_INNER=1 LXC_EXEC="$LXC_EXEC" bash "$0" "$@" +fi + +# The resolver file the sandbox will read is the one the symlink chain ends at, +# so that is the path the stub has to be published at. +RESOLVER_TARGET="$(readlink -f /etc/resolv.conf || true)" +if [ -z "$RESOLVER_TARGET" ] || [ ! -f "$RESOLVER_TARGET" ]; then + echo "SKIP: /etc/resolv.conf does not resolve to a file to publish the stub at." + exit 77 +fi + +WORK_DIR="$(mktemp -d)" +PEER_PID="" +cleanup() { + if [ -n "$PEER_PID" ]; then + kill "$PEER_PID" 2>/dev/null || true + wait "$PEER_PID" 2>/dev/null || true + fi + rm -rf "$WORK_DIR" +} +trap cleanup EXIT + +# Without this the stub address is unreachable and the peer cannot be bound. +ip link set lo up + +# Answers every A query with one fixed record, which is enough to tell a +# resolver that reached the peer from one that reached nothing. AAAA is +# answered empty so a dual-stack lookup falls back to the A record instead of +# waiting out its timeout. +python3 - "$STUB_ADDRESS" "$PROBE_ANSWER" "$WORK_DIR/peer.ready" \ + >"$WORK_DIR/peer.log" 2>&1 <<'PY' & +import socket +import struct +import sys + +address, answer, ready_path = sys.argv[1], sys.argv[2], sys.argv[3] +record = socket.inet_aton(answer) + +with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as server: + server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) + server.bind((address, 53)) + with open(ready_path, "w", encoding="ascii") as ready: + ready.write("ready") + while True: + query, peer = server.recvfrom(4096) + if len(query) < 12: + continue + cursor = 12 + while cursor < len(query) and query[cursor] != 0: + cursor += 1 + query[cursor] + end = cursor + 5 + question = query[12:end] + qtype = struct.unpack("!H", query[end - 4:end - 2])[0] + if qtype == 1: + header = query[:2] + struct.pack("!HHHHH", 0x8180, 1, 1, 0, 0) + body = b"\xc0\x0c" + struct.pack("!HHIH", 1, 1, 60, 4) + record + else: + header = query[:2] + struct.pack("!HHHHH", 0x8180, 1, 0, 0, 0) + body = b"" + server.sendto(header + question + body, peer) +PY +PEER_PID=$! +for _ in $(seq 1 100); do + [ -s "$WORK_DIR/peer.ready" ] && break + if ! kill -0 "$PEER_PID" 2>/dev/null; then + cat "$WORK_DIR/peer.log" + echo "FAIL: the stub resolver exited before it was listening." + exit 1 + fi + sleep 0.1 +done +if [ ! -s "$WORK_DIR/peer.ready" ]; then + cat "$WORK_DIR/peer.log" + echo "FAIL: the stub resolver never signalled readiness." + exit 1 +fi + +printf 'nameserver %s\nsearch %s\noptions edns0 trust-ad\n' \ + "$STUB_ADDRESS" "$SEARCH_DOMAIN" >"$WORK_DIR/resolv.conf" +if ! mount --bind "$WORK_DIR/resolv.conf" "$RESOLVER_TARGET"; then + echo "FAIL: could not publish the stub resolver at $RESOLVER_TARGET." + exit 1 +fi + +# The control. A sandbox that cannot resolve proves nothing unless the peer +# answers the same name from outside it. +if ! HOST_ANSWER="$(timeout 10 getent hosts "$PROBE_NAME" 2>/dev/null)" || + ! grep -Fq "$PROBE_ANSWER" <<<"$HOST_ANSWER"; then + cat "$WORK_DIR/peer.log" + echo "FAIL: $PROBE_NAME does not resolve outside the sandbox either;" + echo " the stub resolver is broken, so a sandbox failure would prove nothing." + exit 1 +fi + +echo "Running Bubblewrap loopback-resolver DNS test..." +RC=0 +OUTPUT="$("$LXC_EXEC" --experimental "$CONFIG" 2>&1)" || RC=$? +if [ "$RC" -ne 0 ]; then + echo "$OUTPUT" + echo "FAIL: the sandbox did not complete its lookup." + exit 1 +fi + +SANDBOX_RESOLV="$(sed -n '/^SANDBOX_RESOLV_BEGIN$/,/^SANDBOX_RESOLV_END$/p' <<<"$OUTPUT")" +if grep -Fq "$STUB_ADDRESS" <<<"$SANDBOX_RESOLV"; then + echo "$OUTPUT" + echo "FAIL: the sandbox still names $STUB_ADDRESS, which is its own empty loopback." + exit 1 +fi +if ! grep -Fq "$SLIRP_FORWARDER" <<<"$SANDBOX_RESOLV"; then + echo "$OUTPUT" + echo "FAIL: the sandbox was not pointed at slirp's forwarder $SLIRP_FORWARDER." + exit 1 +fi +echo "PASS: the sandbox reads slirp's forwarder instead of the host's loopback stub." + +# A resolver that lost the search list resolves a different set of names than +# the host does, which the address assertions above cannot see. +if ! grep -Fq "search $SEARCH_DOMAIN" <<<"$SANDBOX_RESOLV"; then + echo "$OUTPUT" + echo "FAIL: the pinned resolver dropped the host's search list." + exit 1 +fi +echo "PASS: the pinned resolver keeps the host's other directives." + +# The assertion the issue is actually about: a name resolves. +if grep -Fq "MXC_DNS_LOOKUP_FAILED" <<<"$OUTPUT" || + ! grep -Fq "$PROBE_ANSWER" <<<"$OUTPUT"; then + echo "$OUTPUT" + cat "$WORK_DIR/peer.log" + echo "FAIL: $PROBE_NAME did not resolve inside the sandbox, though it resolves outside it." + exit 1 +fi +echo "PASS: a hostname resolves inside the sandbox through slirp's forwarder." +echo "Bubblewrap loopback-resolver DNS test complete."