When Spinifex was first released, we were injecting SSH keys and user data directly from cloud-init. This limited us as it meant anything an application needed at runtime must be populated before the instance boots.
To fix this we need to serve the instance metadata service from the host.
IMDS is a read-only HTTP endpoint at 169.254.169.254 that answers
questions about the instance asking. At boot time it fills network interfaces
and configuration, and at application time it mints temporary credentials. The
key requirement was that stock cloud images could boot exactly how they do on
AWS, meaning no in-guest agent, all the handling must be done outside of the instance.
How AWS does it
On an AWS EC2 instance the host CPUs are not in the network datapath at all.
VPC networking runs on the Nitro card, and the NIC the guest sees is a
virtual function on it. The card already holds the mapping from that
function to the ENI and from the ENI to the instance, so it can easily
identify where the request came from. A request to 169.254.169.254 also never touches the host network.
169.254.169.254 itself
rather than forwarding it.
The card never needs to figure out who is asking. The request arrives on a
virtual function which is tied to only one ENI. Two instances on that server
can both run 10.0.0.5 and it doesn't matter.
We do not have a card. Our datapath is OVS and OVN on the host's own CPU, and the request is a frame on a software bridge beside every other tenant's. We needed a place in that path where a request could only have come from one instance.
Attempt one: answer it on the guest's own switch
169.254.169.254 is a link-local address, and a router will not forward
a packet with a link-local source or destination. Whatever answers the guest has
to sit on the guest's own segment.
We gave each subnet's logical switch an OVN localport claiming
169.254.169.254, sitting beside the guest's own port. The guest
ARPs for the address, the localport answers, and the frame never leaves the
broadcast domain.
A localport exists once in the OVN database, with every chassis creating its own copy, meaning the traffic never leaves the host it lands on. Each host wired a veth pair to its copy and ran a listener on the host end. We then worked out the instance from a pair. The connection arrived on one subnet's veth, which told us the VPC, and the source IP told us which ENI is inside it.
There are three major issues with this design.
The object count was per subnet, on every chassis, whether or not a VM was there. Each subnet materialised a network namespace, a veth pair, a bridge port and a bound listener on every host in the cluster. Twenty-five servers against two thousand subnets is fifty thousand of each object.
The reply needed a routing domain of its own. Binding a socket
to a device pins which device a reply leaves by and supplies no route to get there,
and a single host routing table cannot hold two tenants using the same CIDR. So
the host end of each veth lived in a network namespace per subnet. That meant
creating listeners with an in-process setns, which meant adding
the CAP_SYS_ADMIN capability to the VPC daemon. It also meant the
daemon could hold no filesystem sandbox at all, because in-process
setns and any sandbox directive that forks a mount namespace are
mutually exclusive. This meant stripping ProtectSystem,
ProtectHome, PrivateTmp,
ProtectKernelTunables and ProtectProc off the daemon.
Fixing overlapping tenant CIDRs cost us the hardening on a privileged daemon.
And the guest needed a route. A localport answers only if the
guest treats 169.254.169.254 as on-link and ARPs for it. This is
handled through DHCP option 121, but for static address guests they never ask
DHCP. This meant we had to put it in cloud-init network config instead.
The reason for building IMDS was specifically so we could drop the boot time cloud-init config. But because metadata was only reachable after cloud-init had configured the network, cloud-init could not read user-data from IMDS. This meant every instance kept getting a small disk attached with its user-data, keys and network config on it.
Attempt two: catch it at the tap
The closest thing we have to a Nitro card's virtual function is the tap device. One tap, one ENI, created by the host before the guest boots. The first design asked which ENI sent a request and answered it through a broadcast domain. With a tap, a frame on that tap could only have come from one specific ENI, so we don't need to figure out who requested it.
Each host watches the taps of the VMs it is running. For every primary-ENI
tap it builds a host endpoint of its own. Flows on that tap pull two
destinations aside before anything reaches the integration bridge: 169.254.169.254 for metadata, 169.254.169.253 for DNS. Everything else the guest
sends carries straight on.
A responder binds to each endpoint. It resolves the ENI at that moment and keeps it, then answers whatever arrives, because only one guest's frames can reach that socket.
The request itself does not contain information to identify which instance
it is from. The listener that accepted the connection is the identity,
because a frame only reaches that socket by leaving that tap. Two tenants
can both run 10.0.0.5 and no lookup has to disambiguate them.
We put the flows on a dedicated bridge rather than br-int, as
ovn-controller wipes anything it does not own on every reconnect.
Normal VPC traffic patches straight back to br-int.
This new design solved the three problems from the previous approach.
No route inside the guest. Now that the guest doesn't need to
be configured before it can request IMDS, cloud-init can boot on its standard
Ec2 datasource and take user-data, keys and network config from there.
No namespaces, and the sandbox back. The endpoint sits in the
root namespace, so there is no setns and no
CAP_SYS_ADMIN. The daemon runs under its full systemd sandbox
again.
Bounded cost. Serving state is per local VM rather than per subnet on every host. Now we only need an OVS internal port, a handful of flows, a routing rule and a route, plus the sockets. The whole set is one capture and a host running one VM only needs one.
Close
Spinifex now boots a stock cloud image the way EC2 does. You can launch a stock Ubuntu or Debian image, attach a role, and an SDK inside the instance finds credentials at the address it already looks at.
Take a look at Spinifex on GitHub, and the docs cover single-node, multi-node
and air-gapped installs. If you would rather not install anything yet, the live sandbox gives you a real endpoint in under a minute: launch an instance, attach a role,
and curl 169.254.169.254 from inside it.