Skip to content
Free Spinifex sandbox. Create a sandbox Your Spinifex account is live. Access your console

Self-healing clusters and Terraform plans that stay quiet

A cluster now survives losing a node, with its state replicated everywhere and instances relaunching on their own. Instances take up to eight GPUs, object storage got a sturdier shard layout, and Terraform plans and IAM responses came into line with AWS.

Terraform and OpenTofu: applies that succeed

  • The terraform-aws-modules/eks module applies clean and re-plans with no changes.
  • ebs_block_device and the aws_vpc data source work, and parallel applies no longer lose load balancer targets or tags.
  • Load balancer attributes are accepted in full, and a retried apply no longer duplicates the resource.

Terraform and OpenTofu: plans that report no changes

  • A no-op plan no longer destroys and recreates a working instance.
  • Permanent diffs are cleared on user_data, enclave_options, security group rule tags and ECS availability_zone_rebalancing.
  • ECR and ECS return the fields you set at create time, and most_recent picks the newest AMI.

Identity (IAM and STS): AWS conformance

  • Requests and responses are tested against AWS's published API models for ten services.
  • IAM responses carry the fields AWS returns for each action, with policy documents URL-encoded as AWS does.
  • Managed policy versions, capped at five as in AWS, so editing an aws_iam_policy updates it in place.
  • Trust policies evaluate conditions, and STS rejects session policies and tags instead of ignoring them.

Reliability

  • Instances on a failed node relaunch on a surviving one automatically.
  • Cluster state is replicated to every node, so losing one no longer takes the cluster down.
  • Spinifex deploys on Oracle Cloud Infrastructure (OCI), in beta.

GPU compute

  • Up to eight GPUs per instance, with g5.48xlarge and p4d.24xlarge matching AWS's GPU counts.
  • A new gpu.<count>x<vcpu>c family for smaller edge hosts, such as gpu.8x4c.

Storage (S3)

  • Bucket versioning, user metadata, CopyObject, UploadPartCopy and batch DeleteObjects are available.
  • Keys with special characters are listed by their decoded name, and ETags are derived from content.

Object storage engine (Predastore)

  • Shards spread evenly across nodes, so losing one no longer degrades the whole keyspace.
  • Streaming shard reads and writes with hedged repair, and ranged GETs read only the stripes they touch.
  • Predastore runs standalone with systemd integration.

Networking and security

  • Policy evaluation fails closed, and sts:AssumeRole is gated on the caller's identity policy.
  • Auto-assigned public IPs are released on stop and reassigned on start, as in AWS.
  • Outbound SMTP to public destinations is blocked by default, matching AWS.

RDS arrives, and the firewall comes on by default

The largest month of the year. Managed relational databases went from nothing to two engines, TLS-by-default and drift-free Terraform in three weeks. Underneath it, the install path grew separate WAN, LAN and VPC interfaces and started arming an nftables host firewall, so the cluster's own ports stopped being reachable from anything that could route to a node.

Databases (RDS)

  • Amazon RDS is available: managed PostgreSQL and MariaDB, driven from the AWS CLI, the SDKs, or Terraform's aws_db_instance, across the db.t3, db.m5 and db.r5 families.
  • Every database gets a private endpoint inside your VPC, locked down by security group and never publicly addressable. Connections require TLS by default on both engines.
  • Manual snapshots, automated backups in your window, and restore-to-new-instance, with a new RDS console covering instances, subnet groups, parameter groups and snapshots.

Installation and deployment

  • WAN, LAN and VPC each get their own network interface, so public traffic, internal cluster traffic and VPC overlay traffic no longer share one NIC. A plane left without a NIC folds onto the one above it, so one-, two- and three-NIC servers install from the same model.
  • ext4 joins ZFS as a supported install layout, and the installer assigns whole drives by role so spinifex and predastore each get dedicated disks.
  • scripts/install-node.sh forms a multi-node cluster over SSH in one command, and ISO-installed servers convert into a cluster rather than being reinstalled.

Security

  • A new nftables host firewall restricts cluster ports to known peers, armed by default on ISO installs and opt-in via --firewall=on for package installs, so it never cuts off services already running.
  • IAM policies are evaluated against the actual resource ARNs a request names across every service, so a policy scoped to one resource no longer authorises all of them.
  • Condition blocks on identity policies are enforced instead of being silently discarded.

Storage and quotas

  • Erasure-coded reads are served from parity, so one node down no longer makes objects unreadable, and reads are hedged across shards so a single stalled node does not set the latency of every read.
  • Per-account service quotas cap every metered resource, with an admin surface for per-account overrides.
  • Each storage node can own its data directory, keeping a disk failure to one shard.

Kubernetes and certificates

  • HA control planes spread across distinct Availability Zones first, so a cluster survives losing an AZ.
  • Certificates issue and auto-renew from a tenant Private CA, and re-importing material re-renders every load balancer referencing that ARN.
  • Rebuilt eks-node and ecs-node images, including GPU variants with open-kernel NVIDIA drivers for Blackwell cards.

Multinode goes real: OVN clustering and HA EKS

July was when a Spinifex cluster stopped being several machines and started being one. Software-defined networking moved onto clustered OVSDB RAFT, the EKS control plane spread across hosts and learned to self-heal, and block storage began reclaiming space instead of growing forever.

Networking

  • Multinode software-defined networking via OVN with clustered OVSDB RAFT, plus routed-NAT external mode with two-tier ingress for cleaner north-south traffic.
  • ALB and NLB endpoints resolve by DNS name, and EKS in-cluster DNS resolves from the VPC.

Kubernetes (EKS)

  • Multinode EKS control plane: run the control plane across several hosts, with automatic recovery and continuous reconciliation after node loss.
  • Control-plane disaster recovery — snapshot etcd and restore an entire cluster from a snapshot.
  • Run GPU workloads on EKS, with NVIDIA GPUs exposed to pods through the device plugin, and NVIDIA Blackwell passthrough stable in the EKS and ECS GPU images.

Compute (EC2)

  • EC2 Launch Templates: capture reusable launch configuration and launch straight from a template, in both the API and the console.
  • EC2 Spot Instances are available, and AMIs can be created from volume snapshots and booted directly.
  • Ubuntu NVIDIA GPU node images with automatic GPU-aware AMI selection, and GPU passthrough and MIG partition state visible across the EC2, EKS and ECS console.

Resource tagging

  • Full EC2 tagging — create-tags, delete-tags and describe-tags across instances, volumes, snapshots, AMIs, key pairs, VPCs, gateways, route tables, EIPs and placement groups.
  • Tag at creation with --tag-specifications, and filter describe calls by tag:<key>.
  • IAM tagging for users, roles, policies, instance profiles and OIDC providers, plus tag reads on load balancers, target groups and listeners.

Identity and storage

  • IAM Groups with managed and inline policies, member inheritance, and inline policies on IAM users, editable from the console.
  • Presigned S3 URLs — time-limited GET and PUT links authenticated via SigV4 query-string auth.
  • Block storage reclaims space: unreferenced chunks are garbage-collected and overlapping writes coalesced, so volume growth stays bounded. Enabled by default from v1.14.0.

Observability

  • Full-stack telemetry across the platform: distributed tracing, metrics and structured logs, with per-VM guest metrics.
  • Control-plane, S3 and block-storage logs stream to Elasticsearch via an OTLP bridge.

Kubernetes, containers, and a registry to feed them

Managed Kubernetes landed on a K3s control plane with IAM-authenticated kubectl and IRSA, ECS followed with task definitions and services on EC2 capacity, and ECR closed the loop so workers could pull images without external DNS. Encryption at rest became the default for new installs.

Kubernetes (EKS)

  • A new AWS-compatible managed Kubernetes service on a K3s control plane: full cluster lifecycle, managed node groups whose workers auto-join on boot, and managed add-ons.
  • IAM-authenticated kubectl — aws eks get-token resolves to AccessEntries and access policies, with no aws-auth ConfigMap. IRSA ships with a per-cluster OIDC provider and JWKS discovery.
  • Highly available control plane: three servers across distinct hosts, each cluster in its own managed VPC with private subnets and a NAT gateway for image pulls.
  • Works with stock terraform-aws-eks and eksctl out of the box. Persistent volumes arrive via the aws-ebs-csi-driver addon, and L7 ALB Ingress via aws-load-balancer-controller.

Containers (ECS)

  • Amazon ECS arrives: register task definitions, run tasks, and run long-lived services on EC2 capacity.
  • Pick awsvpc, bridge or host networking per task — awsvpc gives each task its own ENI and VPC IP.
  • Services register with ELBv2 target groups, tasks assume IAM roles through the task credential endpoint, and capacity providers provision EC2 instances on demand.

Container Registry (ECR)

  • A new Elastic Container Registry: create repositories, then push and pull OCI images with docker, crane and skopeo.
  • EKS workers pull from the internal registry without external DNS, for air-gapped clusters.
  • Lifecycle policies expire old images on a background sweep, and immutable tags stop a pushed tag being overwritten onto a different image.

Security

  • At-rest volume encryption works end-to-end and is on by default for new installs, covering encrypted boot volumes, attach and detach, ModifyVolume and snapshots. CMMC Level 1 compliant.
  • Intra-AZ traffic is encrypted by default over IPsec.
  • The web console signs in with short-lived STS session credentials instead of static long-lived keys.

Compute and instance metadata

  • Instances self-configure at boot from instance metadata using cloud-init's standard Ec2 datasource, so stock cloud images boot unmodified and one AMI serves every instance.
  • Each instance gets its own metadata endpoint at 169.254.169.254, serving identity, network and per-MAC interface metadata, with IMDSv2 reported faithfully.
  • NVIDIA MIG GPU slicing, selectable per instance via mig.<profile> types, with GPU inventory in the admin UI.
  • Capacity Reservations: reserve capacity and launch instances directly into a reservation.

Load balancing

  • Network Load Balancers with an L4 data plane on nginx and active per-target health checks, HTTPS listeners with TLS termination, and listener rules routing on host, path, header, method, source IP and query.
  • A new ACM-compatible service for BYO certificates, powering HTTPS listeners.
  • Hosts track real available memory and stop accepting VMs they cannot fit, ending overcommit-driven OOMs.

Security groups start filtering, and FIPS lands

Security group rules stopped being advisory. Ingress and egress reconcile into OVN ACLs with default-deny, which is the difference between a rule you can read back and a rule that drops a packet. FIPS 140-3 cryptography is enforced at startup across every binary, GPU passthrough gained AMD alongside NVIDIA, and a TUI installer replaced hand-assembly of a node.

Security

  • Security groups are enforced end-to-end: ingress and egress rules reconcile into OVN ACLs with default-deny, so they filter traffic rather than describing an intent.
  • FIPS 140-3 cryptography is enforced at startup across all binaries (Go Cryptographic Module v1.0.0, CMVP cert #5247).
  • TLS 1.3 minimum across the stack with hybrid post-quantum key exchange (X25519 + ML-KEM), and Raft cluster traffic TLS-wrapped.

Compute and GPU

  • NVIDIA VFIO GPU passthrough with auto-discovered g5.* types and a driver-preinstalled Ubuntu GPU AMI, plus spx admin gpu status|enable|disable with live config reload.
  • AMD support (MI350X) alongside NVIDIA, and multi-GPU instances with the new g7e.12xlarge.
  • UEFI boot is supported and is the default for new instances. Rocky Linux, RHEL, Ubuntu 26.04 and Debian 13 images joined the catalog.
  • Load balancers launch via direct-boot QEMU microvm with no system AMI import — boot time dropped 12×, from 3.2 s to 264 ms.

Installation

  • A new TUI-based ISO installer with headless autoinstall.toml mode, GRUB-driven disk selection, multi-NIC WAN/LAN bridge setup and a dedicated spinifex login user.
  • spx admin init/join splits listen from advertise addresses, provisions br-mgmt over OVS, and supports formation join tokens.
  • Two-node distributed predastore mode using RS(1,1) mirroring, which previously required three nodes.

Networking and reliability

  • New tenant accounts auto-provision a default VPC and internet gateway, so RunInstances works out of the box on a fresh tenant.
  • The daemon survives NATS outages, serving locally when the cluster control plane is unreachable, with local-first instance state and a read-only /local/* API for operators.
  • StartInstances routes back to the node that last ran the instance, where its volumes live, with automatic fallback on node-down or capacity exhaustion.

Hardening the perimeter after the 1.0 launch

NATS authentication became mandatory, `InsecureSkipVerify` was removed, services picked up systemd hardening directives and privilege separation, and the AWS gateway took shape as its own component. Load balancing made its first appearance.

Security

  • NATS authentication is enforced, InsecureSkipVerify is removed, and TLS is enabled on the cluster manager API.
  • Privilege separation and systemd hardening directives across all service units.

Load balancing (ELBv2)

  • The initial ELBv2 implementation, with an attributes API, DescribeTags, and a load balancer frontend.

Platform

  • The AWS gateway (awsgw) takes shape as its own component, alongside a webserver proxy and a migration framework.
  • A new nginx-ALB Terraform example with private-IP instances.

Spinifex 1.0 — cloud-native software, no cloud required

Spinifex 1.0 released on the 31st of March. The pitch has not changed since: replicate a hyperscale cloud environment on compute you own, so cloud-native applications and AI workloads run anywhere, independent of centralised infrastructure.

The 1.0 release

  • The basics, working end to end: launch and terminate instances, create and attach volumes, take snapshots, put and get objects, and carve out a VPC to run it all in.
  • Install with a single command: curl -fsSL https://install.mulgadc.com | bash, then spx admin init and systemctl start spinifex.target.
  • Point the AWS CLI at it with a profile and start making EC2 calls.
  • The first release we were happy to put a 1.0 on — enough of the surface was real that you could run something on it rather than evaluate it.