Pinned Loading
Repositories
- scheduler-vs-more-gpus Public
Turning scattered GPUs into working AI clusters: capacity accounting, an allocation-policy simulator, control-plane case studies, and a TPU parallel hypothesis. Slurm, Kubernetes, multi-cloud. Reproducible.
- research Public
DIMAGGI AI infrastructure research — index of the usable-capacity series and installable standards. https://dimaggi-ai.github.io/research/
- airan-neocloud-resiliency Public
Monte Carlo resilience of the AI-RAN last mile, priced in nines: a single-fiber edge site is ~2.7 nines; multi-transport bonding helps only the tenant classes; the generator — not a third link — buys the decisive nine.
- edge-continuum-placement Public
Where AI workloads belong on the operator's footprint — tower, hub, metro PoP, central AI factory. A gate-cascade placement engine (fabric, power, latency, gravity): centralize by default, edge when forced. The tower is never the recommended tier.
- compute-power-placement Public
When does moving an AI training job to cheaper/available power beat staying? A move-vs-stay decision model in GPU-hours and dollars. Finding: schedule against scarcity, don't arbitrage the spread.
- governed-autonomy Public
Architecture for an autonomous control plane over GPU clusters and networks: separated planes, a referee, an autonomy ladder earned by chaos experiments, and a sourced latency hierarchy — from human intent to nanosecond in-ASIC reflexes.
- reliability-economics Public
Which GPU-cluster recovery policy wins in which failure regime? A reproducible simulator (ETTR/MTTF/MTTR + dollars) and a policy phase diagram, validated against Meta's published reliability numbers. Spares vs elastic vs checkpointing.
- ai-cluster-chaos-fidelity Public
A machine-checkable standard for AI-cluster chaos engineering: which fault injection tests which layer. Fidelity linter + spec schema + 20-experiment catalog. A pod kill is not an XID; tc/netem is not an InfiniBand flap.
- network-vs-more-gpus Public
When does network, reliability, or recovery investment beat buying more GPUs? Accounting framework, substitution metric, and decision maps for AI training infrastructure. Paper + reproducible model.
- cooling-pue-ladder Public
The cooling ladder, executable: density gates, PUE-to-megawatts capacity math, ride-through seconds, and the two-fleets explanation of the PUE plateau
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…