Skip to content
Free Spinifex sandbox. Create a sandbox Your Spinifex account is live. Access your console
All posts
Engineering · 6 min read

We Built Our Metadata Service Twice

How we serve instance metadata without a Nitro card

An instance needs somewhere to fetch credentials. We served 169.254.169.254 on the guest's subnet switch, then at the instance's tap. Here is what each cost.

JS

Julian Sommer

Systems Engineer

When Spinifex was first released, we were injecting SSH keys and user data directly from cloud-init. This limited us as it meant anything an application needed at runtime must be populated before the instance boots.

To fix this we need to serve the instance metadata service from the host. IMDS is a read-only HTTP endpoint at 169.254.169.254 that answers questions about the instance asking. At boot time it fills network interfaces and configuration, and at application time it mints temporary credentials. The key requirement was that stock cloud images could boot exactly how they do on AWS, meaning no in-guest agent, all the handling must be done outside of the instance.

How AWS does it

On an AWS EC2 instance the host CPUs are not in the network datapath at all. VPC networking runs on the Nitro card, and the NIC the guest sees is a virtual function on it. The card already holds the mapping from that function to the ENI and from the ENI to the instance, so it can easily identify where the request came from. A request to 169.254.169.254 also never touches the host network.

The Nitro metadata path: the guest's NIC is a virtual function on the Nitro card, the card already holds that function's ENI and instance binding for VPC encapsulation, and it answers 169.254.169.254 itself rather than forwarding it.
The Nitro metadata path. The card holds the virtual function's ENI and instance binding already, so it answers 169.254.169.254 itself rather than forwarding it.

The card never needs to figure out who is asking. The request arrives on a virtual function which is tied to only one ENI. Two instances on that server can both run 10.0.0.5 and it doesn't matter.

We do not have a card. Our datapath is OVS and OVN on the host's own CPU, and the request is a frame on a software bridge beside every other tenant's. We needed a place in that path where a request could only have come from one instance.

Attempt one: answer it on the guest's own switch

169.254.169.254 is a link-local address, and a router will not forward a packet with a link-local source or destination. Whatever answers the guest has to sit on the guest's own segment.

We gave each subnet's logical switch an OVN localport claiming 169.254.169.254, sitting beside the guest's own port. The guest ARPs for the address, the localport answers, and the frame never leaves the broadcast domain.

A localport exists once in the OVN database, with every chassis creating its own copy, meaning the traffic never leaves the host it lands on. Each host wired a veth pair to its copy and ran a listener on the host end. We then worked out the instance from a pair. The connection arrived on one subnet's veth, which told us the VPC, and the source IP told us which ENI is inside it.

The subnet-switch datapath: the localport claiming 169.254.169.254 sits on the guest's own logical switch beside the guest port, one L2 hop away, served by a host veth in a per-subnet network namespace, with the VPC logical router never in the path.
The subnet-switch datapath. The localport sits on the guest's own logical switch, one L2 hop away, served by a host veth in a per-subnet network namespace.

There are three major issues with this design.

The object count was per subnet, on every chassis, whether or not a VM was there. Each subnet materialised a network namespace, a veth pair, a bridge port and a bound listener on every host in the cluster. Twenty-five servers against two thousand subnets is fifty thousand of each object.

The reply needed a routing domain of its own. Binding a socket to a device pins which device a reply leaves by and supplies no route to get there, and a single host routing table cannot hold two tenants using the same CIDR. So the host end of each veth lived in a network namespace per subnet. That meant creating listeners with an in-process setns, which meant adding the CAP_SYS_ADMIN capability to the VPC daemon. It also meant the daemon could hold no filesystem sandbox at all, because in-process setns and any sandbox directive that forks a mount namespace are mutually exclusive. This meant stripping ProtectSystem, ProtectHome, PrivateTmp, ProtectKernelTunables and ProtectProc off the daemon. Fixing overlapping tenant CIDRs cost us the hardening on a privileged daemon.

And the guest needed a route. A localport answers only if the guest treats 169.254.169.254 as on-link and ARPs for it. This is handled through DHCP option 121, but for static address guests they never ask DHCP. This meant we had to put it in cloud-init network config instead.

The reason for building IMDS was specifically so we could drop the boot time cloud-init config. But because metadata was only reachable after cloud-init had configured the network, cloud-init could not read user-data from IMDS. This meant every instance kept getting a small disk attached with its user-data, keys and network config on it.

Attempt two: catch it at the tap

The closest thing we have to a Nitro card's virtual function is the tap device. One tap, one ENI, created by the host before the guest boots. The first design asked which ENI sent a request and answered it through a broadcast domain. With a tap, a frame on that tap could only have come from one specific ENI, so we don't need to figure out who requested it.

Each host watches the taps of the VMs it is running. For every primary-ENI tap it builds a host endpoint of its own. Flows on that tap pull two destinations aside before anything reaches the integration bridge: 169.254.169.254 for metadata, 169.254.169.253 for DNS. Everything else the guest sends carries straight on.

A responder binds to each endpoint. It resolves the ENI at that moment and keeps it, then answers whatever arrives, because only one guest's frames can reach that socket.

The per-tap datapath: two guests holding the same address each reach their own endpoint on a dedicated bridge, which demultiplexes by destination to that tap's metadata responder and DNS shim, while all other traffic patches through to the OVN integration bridge.
The per-tap datapath. Two guests holding the same address each reach their own endpoint, demultiplexed by destination to that tap's metadata responder and DNS shim.

The request itself does not contain information to identify which instance it is from. The listener that accepted the connection is the identity, because a frame only reaches that socket by leaving that tap. Two tenants can both run 10.0.0.5 and no lookup has to disambiguate them.

We put the flows on a dedicated bridge rather than br-int, as ovn-controller wipes anything it does not own on every reconnect. Normal VPC traffic patches straight back to br-int.

This new design solved the three problems from the previous approach.

No route inside the guest. Now that the guest doesn't need to be configured before it can request IMDS, cloud-init can boot on its standard Ec2 datasource and take user-data, keys and network config from there.

No namespaces, and the sandbox back. The endpoint sits in the root namespace, so there is no setns and no CAP_SYS_ADMIN. The daemon runs under its full systemd sandbox again.

Bounded cost. Serving state is per local VM rather than per subnet on every host. Now we only need an OVS internal port, a handful of flows, a routing rule and a route, plus the sockets. The whole set is one capture and a host running one VM only needs one.

Close

Spinifex now boots a stock cloud image the way EC2 does. You can launch a stock Ubuntu or Debian image, attach a role, and an SDK inside the instance finds credentials at the address it already looks at.

Take a look at Spinifex on GitHub, and the docs cover single-node, multi-node and air-gapped installs. If you would rather not install anything yet, the live sandbox gives you a real endpoint in under a minute: launch an instance, attach a role, and curl 169.254.169.254 from inside it.

#IMDS#Networking#OVN#Multi-Tenancy#Spinifex