Skip to content
Free Spinifex sandbox. Create a sandbox Your Spinifex account is live. Access your console
All posts
Engineering · 14 min read

Migrating from MinIO to Predastore

A step-by-step guide for single-node and distributed deployments

MinIO is archived and no longer patched. This guide walks through standing up Predastore, recreating users and policies, and copying every bucket across.

BM

Bryn Mailer

Senior Systems Engineer

If you're reading this, you're probably already aware of MinIO's fall from grace. The open source community's favorite object storage system was officially archived back in April (2026, for future readers). The impending deficit of security patches and bug fixes has since sent many sysadmins looking for a maintained alternative.

If you're one of the affected individuals, I'd like to take the opportunity to introduce you to Predastore!

My team and I have been working on this project for around a year and a half now. Long before the crescendo of the MinIO saga. The original motivation was to build an object storage system capable of operating in resource constrained edge environments with an unreliable network fabric.

Although there is still plenty of work to be done to soundly meet that goal, the project in its current state happens to fit the more standard use cases (that MinIO excelled in) rather well!

By the end of this article, you will hopefully have a clear understanding of the process of migrating from a MinIO deployment to a brand new Predastore cluster.

An architectural overview

Diagram of a simple Predastore cluster: S3 clients connect over HTTPS to the gate node on each of three hosts. Each host has its own IP address and runs a gate, meta and blob node on separate ports, and each blob node writes to its own disk.

Note: Everything in this section is subject to change. Predastore is a fairly young project, and certain design decisions have not been set in stone yet. For example, I'm not particularly sold on the "blob" terminology yet. If anyone has any better ideas, please raise a GitHub issue! It will make my day.

In order to start building a mental model of how MinIO concepts map to Predastore, a rudimentary of understanding of Predastore itself is required. This will be rather brief. If you want the meatier details, check out the design doc located here.

A Predastore cluster is organized into two levels; hosts, and nodes. A host is just a single Predastore process, invoked via s3d. Each host can contain any number of nodes, where each node has a role that is one of:

  • gate - Exposes the S3 HTTP API. There may be at most one gate node per host process.
  • meta - A Raft replica holding bucket and object records, including where every object's shards live.
  • blob - A key-value store that writes encrypted, erasure-coded shards to append-only segment files. The directory to which a blob node reads/writes is set using the data_dir field in the config file.

Each host gets its own IP address, and each node within a single host binds to a separate port under the host IP address. This ensures that every node is individually addressable.

A cluster must contain at least one of each type of node. Additionally, it is highly recommended that each blob node is pointed at a separate disk to ensure proper isolation.

Before you begin: what Predastore doesn't currently do

MinIO was a mature project, with many years spent adding support for all manner of features provided by S3. Predastore is rather young as of yet, and while we support the main body of features required by most deployments, there are a number of things still on the TODO list.

Before proceeding with the following steps, take a moment to skim the S3 API coverage page here to ensure Predastore isn't missing an operation your use case depends on.

If something you're after is absent, please notify my team and me by raising a GitHub issue. We're pretty responsive. Depending on the request, there's a good chance we'll be able to slip it into a minor release.

1. Pick a mode: standalone or under Spinifex

The first step in the process is to decide whether your new Predastore cluster is going to be deployed standalone, or via Spinifex. This is usually a fairly simple choice, depending on your desired authentication structure. The general capabilities of each method:

  • Standalone - A fixed set of service accounts, one tenant, no external dependencies. Manual distribution of keys and certificates to nodes during setup.
  • Via Spinifex - IAM users/groups/roles, STS, multiple tenants. Automatic distribution of keys and certificates to member nodes.

I highly recommend taking a moment to peruse the "Standalone or Spinifex?" section of Predastore's README for guidance on which path to take.

2. Stand up Predastore

The exact installation steps you should follow depend on the structure of your pre-existing MinIO cluster. Use the tabs below to show the steps that match your deployment topology and chosen mode of installation.

Single-Node Single/Multi-Drive - Standalone

This is the simplest path, and the closest equivalent to a single MinIO server. You'll need a Linux machine running systemd, along with Go (1.27 or newer), git, make, and openssl in order to build Predastore from source. Most distro packages of Go are a few versions behind, so you may need to grab it from go.dev. You'll also want the AWS CLI for checking Predastore once it's up.

Build and install

First up, clone the repository and build the s3d binary:

git clone --branch v1.21.0 https://github.com/mulgadc/predastore.git
cd predastore
make build

Note: This guide was written against v1.21.0. The install tooling used below (make install, predastore-keygen, and the example config) first shipped in v1.20.0, so don't go any older than that.

Install it, along with its systemd service:

sudo make install

Then create the dedicated predastore user, and the config and data directories (/etc/predastore and /var/lib/predastore respectively):

sudo systemd-sysusers && sudo systemd-tmpfiles --create

Share MinIO's drives

There's no need for new drives. Predastore can live on the ones MinIO is already using, in a directory of its own alongside MinIO's data, until you're done with MinIO in step 5. Each drive needs roughly as much free space again as MinIO is using on it, plus some headroom for re-syncs, which df -h will tell you.

The steps below assume MinIO's drives are mounted at /mnt/disk1, /mnt/disk2, and so on. If they're pooled with RAID or ZFS, treat the pool as a single drive.

For each drive, create a directory for Predastore. The leading . matters, as MinIO treats any other top-level directory on its drives as a bucket, and will delete it along with that bucket.

sudo install -d -o predastore -g predastore -m 0700 /mnt/disk1/.predastore

The service isn't permitted to write anywhere outside /var/lib/predastore, so bind-mount each directory into it. The extra options make sure the service waits for the mount, and refuses to start without it:

sudo install -d -m 0755 /var/lib/predastore/disk1
echo '/mnt/disk1/.predastore /var/lib/predastore/disk1 none bind,nofail,x-systemd.requires-mounts-for=/mnt/disk1,x-systemd.required-by=predastore.service,x-systemd.before=predastore.service 0 0' | sudo tee -a /etc/fstab

Once every drive has its line in /etc/fstab, mount them:

sudo systemctl daemon-reload
sudo mount -a

Note: Leave /var/lib/predastore/disk1 and friends owned by root. If a mount ever goes missing, that's what stops Predastore from quietly writing to the system disk in its place.

Match your MinIO setup

Predastore ships with an example config that already describes a working single-drive host, so there isn't a great deal to change. Copy it into place, and open it up in your editor of choice:

sudo cp /etc/predastore/predastore.toml.example /etc/predastore/predastore.toml
sudo $EDITOR /etc/predastore/predastore.toml

For the purposes of the migration, there are only two fields you need to set. region needs to match whatever your MinIO clients are configured with, as S3 clients sign every request for a specific region. [[auth]] holds the access key and secret your applications will use.

Keep in mind that these are root credentials, with full access to every bucket in the cluster (as mentioned in the standalone vs Spinifex discussion).

region = "us-east-1"

[[auth]]
access_key_id     = "<access-key>"
secret_access_key = "<secret-key>"
account_id        = "100000000001"

The example config also includes an example-bucket entry, which you can go ahead and delete.

Next, point the blob nodes at the drives. If MinIO only had one drive, set data_dir = "/var/lib/predastore/disk1" on the example's blob node, and the config is done.

Otherwise, replace the example's single blob node with one blob node per drive. Give each one its own id and port, and point its own data_dir at the drive's bind mount.

Leave the host's data_dir as it is, since the meta node doesn't name a directory of its own and keeps its state under /var/lib/predastore on the system disk. With three drives, the [[host]] block ends up looking like this:

[[host]]
id = 1
addr = "127.0.0.1"
admin_port = 9099
data_dir = "/var/lib/predastore"
tls_cert = "/etc/predastore/server.pem"
tls_key = "/etc/predastore/server.key"
encryption_key = "/etc/predastore/master.key"

  [[host.node]]
  id = 1
  role = "gate"
  port = 8443
  bind_addr = "0.0.0.0"

  [[host.node]]
  id = 2
  role = "meta"
  port = 6660

  [[host.node]]
  id = 3
  role = "blob"
  port = 9991
  data_dir = "/var/lib/predastore/disk1"

  [[host.node]]
  id = 4
  role = "blob"
  port = 9992
  data_dir = "/var/lib/predastore/disk2"

  [[host.node]]
  id = 5
  role = "blob"
  port = 9993
  data_dir = "/var/lib/predastore/disk3"

Finally, set [rs] to match. Each object is split into data shards plus parity shards, each on a different drive, so parity is the number of drives you can lose, and data + parity can't exceed the number of drives.

RS(2,1) across three drives will survive losing any one of them (for two drives, use RS(1,1)). Your data survives, but the service won't start while a drive's bind mount is missing, so mount the replacement drive at the same path, recreate its .predastore directory, and run sudo mount -a before restarting. [rs] can't be changed once objects have been written, so make sure you're happy with it before moving any data across.

The configuration reference covers the details.

[rs]
data   = 2
parity = 1

Generate keys and a certificate

Generate the key used to encrypt data at rest, along with a self-signed TLS certificate. PREDA_SAN should list every hostname and IP address your clients will use to reach Predastore, otherwise they'll refuse to trust the certificate when you come to copy data across. It replaces the default list rather than adding to it, so include DNS:localhost and IP:127.0.0.1 if you'll also be connecting from the machine itself.

sudo PREDA_SAN="DNS:s3.example.com,IP:192.0.2.10" predastore-keygen /etc/predastore

Note: Make sure to back up /etc/predastore/master.key somewhere separate from your data backups. Without it your data is unreadable, but storing it alongside the data backups defeats the purpose of encrypting at all.

Start the service

Enable and start the service, then follow its logs to make sure it comes up without any errors:

sudo systemctl daemon-reload
sudo systemctl enable --now predastore
sudo journalctl -u predastore -f

Note: A few warnings while the nodes start up are normal, such as the gate reporting it isn't ready yet. It's errors, and the service restarting, that you're looking out for.

Check it's up

With the service running, it's worth confirming that Predastore is reachable, and that your client and cluster agree on the region.

The certificate is only readable by the predastore user, so copy it somewhere your own user can read (and over to any other machine you'll be running clients from). The rest of this guide uses this copy:

sudo install -m 0644 -o "$USER" /etc/predastore/server.pem ~/predastore-ca.pem

Point the AWS CLI at the credentials from your [[auth]] entry, and at the certificate:

export AWS_ACCESS_KEY_ID=<access-key>
export AWS_SECRET_ACCESS_KEY=<secret-key>
export AWS_CA_BUNDLE=~/predastore-ca.pem

Then create, list, and remove a test bucket. Listing buckets alone won't catch a region mismatch, which is why the check creates one.

aws s3 mb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 ls s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 rb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1

If all three commands complete without any errors, your Predastore cluster is ready for data. If you see AuthorizationHeaderMalformed ... incorrect region, the region your client is signing for doesn't match the one Predastore was set up with.

Multi-Node Multi-Drive - Standalone

This path is the equivalent of a distributed MinIO deployment. Every machine runs the same config file, and shares a single TLS certificate and encryption key.

The steps below assume three machines at 10.0.0.1, 10.0.0.2, and 10.0.0.3, but the process is the same for any number of them. You'll need Go (1.27 or newer), git, make, and openssl on each machine, as well as the AWS CLI on at least one of them for checking the cluster once it's up.

Build and install

On every machine, clone the repository and build the s3d binary:

git clone --branch v1.21.0 https://github.com/mulgadc/predastore.git
cd predastore
make build

Note: This guide was written against v1.21.0. The install tooling used below (make install, predastore-keygen, and the example config) first shipped in v1.20.0, so don't go any older than that.

Install it, along with its systemd service:

sudo make install

Create the dedicated predastore user, and the config and data directories:

sudo systemd-sysusers && sudo systemd-tmpfiles --create

make install also ships some kernel network settings. Hosts on separate machines talk to each other over QUIC, and need larger socket buffers than most distros provide by default. Without them, writes will start failing under load, with no obvious indication as to why. Apply them:

sudo sysctl --system

Share MinIO's drives

There's no need for new drives. Predastore can live on the ones MinIO is already using, in a directory of its own alongside MinIO's data, until you're done with MinIO in step 5. Each drive needs roughly as much free space again as MinIO is using on it, plus some headroom for re-syncs, which df -h will tell you.

The steps below assume each machine's MinIO drives are mounted at /mnt/disk1, /mnt/disk2, and so on. If they're pooled with RAID or ZFS, treat the pool as a single drive. The pool covers a failed drive, and Predastore covers a failed machine.

On every machine, create a directory for Predastore on each drive. The leading . matters, as MinIO treats any other top-level directory on its drives as a bucket, and will delete it along with that bucket.

sudo install -d -o predastore -g predastore -m 0700 /mnt/disk1/.predastore

The service isn't permitted to write anywhere outside /var/lib/predastore, so bind-mount each directory into it. The extra options make sure the service waits for the mount, and refuses to start without it:

sudo install -d -m 0755 /var/lib/predastore/disk1
echo '/mnt/disk1/.predastore /var/lib/predastore/disk1 none bind,nofail,x-systemd.requires-mounts-for=/mnt/disk1,x-systemd.required-by=predastore.service,x-systemd.before=predastore.service 0 0' | sudo tee -a /etc/fstab

Once every drive has its line in /etc/fstab, mount them:

sudo systemctl daemon-reload
sudo mount -a

Note: Leave /var/lib/predastore/disk1 and friends owned by root. If a mount ever goes missing, that's what stops Predastore from quietly writing to the system disk in its place.

Describe the cluster

A single config file describes the entire cluster, with one [[host]] block per machine. Each host gets its own address, its data directory, key, and certificate paths, and its own gate, meta, and blob nodes. Node ids must be unique across the whole file, not just within a host.

region needs to match whatever your MinIO clients are configured with, as S3 clients sign every request for a specific region. [[auth]] holds the access key and secret your applications will use.

Keep in mind that these are root credentials, with full access to every bucket in the cluster (as mentioned in the standalone vs Spinifex discussion).

[rs] determines how objects are spread across the machines. Each object is split into data shards plus parity shards, each on a different blob node. RS(2,1) across three machines will survive losing any one of them.

Check out the configuration reference for the full format.

version = 1
region  = "us-east-1"

[rs]
data   = 2
parity = 1

[[host]]
id   = 1
addr = "10.0.0.1"
admin_port = 9099
data_dir = "/var/lib/predastore"
tls_cert = "/etc/predastore/server.pem"
tls_key = "/etc/predastore/server.key"
encryption_key = "/etc/predastore/master.key"

  [[host.node]]
  id = 1
  role = "gate"
  port = 8443
  bind_addr = "0.0.0.0"

  [[host.node]]
  id = 2
  role = "meta"
  port = 6660

  [[host.node]]
  id = 3
  role = "blob"
  port = 9991

[[host]]
id   = 2
addr = "10.0.0.2"
admin_port = 9099
data_dir = "/var/lib/predastore"
tls_cert = "/etc/predastore/server.pem"
tls_key = "/etc/predastore/server.key"
encryption_key = "/etc/predastore/master.key"

  [[host.node]]
  id = 4
  role = "gate"
  port = 8443
  bind_addr = "0.0.0.0"

  [[host.node]]
  id = 5
  role = "meta"
  port = 6660

  [[host.node]]
  id = 6
  role = "blob"
  port = 9991

[[host]]
id   = 3
addr = "10.0.0.3"
admin_port = 9099
data_dir = "/var/lib/predastore"
tls_cert = "/etc/predastore/server.pem"
tls_key = "/etc/predastore/server.key"
encryption_key = "/etc/predastore/master.key"

  [[host.node]]
  id = 7
  role = "gate"
  port = 8443
  bind_addr = "0.0.0.0"

  [[host.node]]
  id = 8
  role = "meta"
  port = 6660

  [[host.node]]
  id = 9
  role = "blob"
  port = 9991

[[auth]]
access_key_id     = "<access-key>"
secret_access_key = "<secret-key>"
account_id        = "100000000001"

Save it as /etc/predastore/predastore.toml on every machine.

Since every machine shares the same config file, each one needs telling which [[host]] block describes it. That lives in /etc/predastore/predastore.env. On each machine, copy the example into place, and open it up:

sudo cp /etc/predastore/predastore.env.example /etc/predastore/predastore.env
sudo $EDITOR /etc/predastore/predastore.env

Set PREDA_HOST_ID to the id of the [[host]] block that describes this machine, so 1 on 10.0.0.1, 2 on 10.0.0.2, and so on. The rest of the file can be left as it is.

PREDA_HOST_ID=1

Next, point the blob nodes at the drives. If each machine only has one drive, add data_dir = "/var/lib/predastore/disk1" to every host's blob node, and the config is done.

Otherwise, give each host one blob node per drive, each with its own id, port, and data_dir pointing at the drive's bind mount. Leave each host's data_dir as it is, since the meta node keeps its state there.

With two drives per machine, host 1 looks like this, and hosts 2 and 3 follow the same pattern with node ids 5 to 8 and 9 to 12:

[[host]]
id   = 1
addr = "10.0.0.1"
admin_port = 9099
data_dir = "/var/lib/predastore"
tls_cert = "/etc/predastore/server.pem"
tls_key = "/etc/predastore/server.key"
encryption_key = "/etc/predastore/master.key"

  [[host.node]]
  id = 1
  role = "gate"
  port = 8443
  bind_addr = "0.0.0.0"

  [[host.node]]
  id = 2
  role = "meta"
  port = 6660

  [[host.node]]
  id = 3
  role = "blob"
  port = 9991
  data_dir = "/var/lib/predastore/disk1"

  [[host.node]]
  id = 4
  role = "blob"
  port = 9992
  data_dir = "/var/lib/predastore/disk2"

Finally, set [rs] to match, which needs some care. Predastore spreads an object's shards across distinct blob nodes, but doesn't take into account which machine each blob node lives on, so two shards of the same object can land on two drives in the same machine.

To survive losing a whole machine, parity needs to be at least the number of drives per machine, and data + parity can't exceed the total number of drives. For three machines with two drives each, RS(2,2) survives losing any one machine. [rs] can't be changed once objects have been written, so make sure you're happy with it before moving any data across.

[rs]
data   = 2
parity = 2

Generate keys and a certificate

Every machine in the cluster needs to share the same TLS certificate and the same encryption key. This one is a bit of a trap; machines with their own certificates never manage to form a cluster.

So rather than running predastore-keygen on each machine straight away, generate the files once on a single machine and distribute them yourself.

On one machine, generate a self-signed certificate. Make sure subjectAltName lists every machine's address, as well as the hostname your clients will use:

sudo openssl req -x509 -newkey rsa:2048 -nodes -days 825 -subj '/CN=predastore' \
  -addext "subjectAltName=DNS:s3.example.com,IP:10.0.0.1,IP:10.0.0.2,IP:10.0.0.3" \
  -keyout /etc/predastore/server.key -out /etc/predastore/server.pem

On the same machine, generate the encryption key:

sudo sh -c 'umask 0177 && openssl rand -out /etc/predastore/master.key 32'

Copy server.pem, server.key and master.key into /etc/predastore on every other machine, using something that preserves file permissions, such as scp -p. Predastore refuses to start if either key is readable by anyone other than its owner, so if in doubt, tighten them up on each machine once they've landed:

sudo chmod 0600 /etc/predastore/master.key /etc/predastore/server.key

Then, on every machine, run predastore-keygen. It keeps the copied files, and sets their ownership:

sudo predastore-keygen /etc/predastore

Still on every machine, add the certificate to the system trust store, so each host trusts the others:

sudo cp /etc/predastore/server.pem /usr/local/share/ca-certificates/predastore.crt
sudo update-ca-certificates

Note: The trust store commands above are for Debian and Ubuntu. For RHEL and Fedora, see the Standalone TLS Trust section of the README. As with the single node setup, make sure to back up master.key somewhere separate from your data backups.

Start the service

On every machine, enable and start the service, then follow its logs:

sudo systemctl daemon-reload
sudo systemctl enable --now predastore
sudo journalctl -u predastore -f

Until every machine is up, the ones that started first will log errors about not being able to reach the others, which is expected. Once the service is running on every machine, the hosts will find each other and elect a leader, and those errors should stop. From then on, it's errors and the service restarting that you're looking out for.

Check it's up

With the cluster running, it's worth confirming that Predastore is reachable, and that your client and cluster agree on the region. Run the following from any one of the machines.

s3.example.com should point at every machine's gate, not just one, either with a DNS record per machine or a load balancer in front of them. Clients only talk to the gates that name resolves to, so if it names a single machine, losing that machine takes Predastore offline for your applications, even though the rest of the cluster carries on.

The certificate is only readable by the predastore user, so copy it somewhere your own user can read (and over to any other machine you'll be running clients from). The rest of this guide uses this copy:

sudo install -m 0644 -o "$USER" /etc/predastore/server.pem ~/predastore-ca.pem

Point the AWS CLI at the credentials from your [[auth]] entry, and at the certificate:

export AWS_ACCESS_KEY_ID=<access-key>
export AWS_SECRET_ACCESS_KEY=<secret-key>
export AWS_CA_BUNDLE=~/predastore-ca.pem

Then create, list, and remove a test bucket. Listing buckets alone won't catch a region mismatch, which is why the check creates one.

aws s3 mb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 ls s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 rb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1

If all three commands complete without any errors, your Predastore cluster is ready for data. If you see AuthorizationHeaderMalformed ... incorrect region, the region your client is signing for doesn't match the one Predastore was set up with.

Single-Node Single/Multi-Drive - Via Spinifex

A caveat before you start: Spinifex currently keeps all of Predastore's data on a single drive. While MinIO is still running, that means picking one of MinIO's drives to hold everything, which limits Predastore to that drive's capacity, with no protection against losing it. If MinIO's drives are already pooled with RAID or ZFS, the pool counts as one drive and keeps its redundancy. Support for spreading Predastore across multiple drives under Spinifex is on its way. Until then, if you need to survive a failed drive and can live without IAM, the standalone path is the better fit.

If you've gone with Spinifex, Predastore comes along for the ride. It's installed and configured as part of Spinifex, so there's no need to set it up separately.

On a single node, Spinifex runs Predastore with one blob node and RS(1,0), which means a single copy of each object with no parity.

Run the installer

Before installing, make sure the machine has a Linux bridge configured for VM networking. The bridge setup guide walks through setting one up.

Run the installer:

curl -fsSL https://install.mulgadc.com | bash

Then set up OVN networking. --management also starts OVN's central services, which a single node needs to run itself:

sudo /usr/local/share/spinifex/setup-ovn.sh --management

Initialize the node

Initializing the node generates Predastore's config, keys, and certificates for you. Just like the standalone setup, --region needs to match your MinIO clients, as Spinifex defaults to ap-southeast-2.

sudo spx admin init --node node1 --nodes 1 --region us-east-1

Share one of MinIO's drives

Predastore can live on one of the drives MinIO is already using, in a directory of its own alongside MinIO's data, until you're done with MinIO in step 5. Pick the drive with the most free space, as it needs enough room for a full copy of your data, plus some headroom for re-syncs.

The steps below assume that drive is mounted at /mnt/disk1. Create a directory for Predastore on it. The leading . matters, as MinIO treats any other top-level directory on its drives as a bucket, and will delete it along with that bucket.

sudo install -d -o spinifex-storage -g spinifex -m 0700 /mnt/disk1/.spinifex-predastore

Predastore's service isn't permitted to write anywhere outside /var/lib/spinifex/predastore, so bind-mount the directory over it. The extra options make sure the service waits for the mount, and won't write to the system disk without it:

sudo chown root:root /var/lib/spinifex/predastore
sudo chmod 0755 /var/lib/spinifex/predastore
echo '/mnt/disk1/.spinifex-predastore /var/lib/spinifex/predastore none bind,nofail,x-systemd.requires-mounts-for=/mnt/disk1,x-systemd.required-by=spinifex-predastore.service,x-systemd.before=spinifex-predastore.service 0 0' | sudo tee -a /etc/fstab
sudo systemctl daemon-reload
sudo mount -a

Start the services

Start Spinifex's services, Predastore included:

sudo systemctl start spinifex.target

Spinifex's certificate only covers the machine's own hostname and addresses. If your clients will reach Predastore by any other name, such as s3.example.com, add it to the certificate, and restart the services to pick it up:

sudo spx admin cert renew --extra-dns s3.example.com
sudo systemctl restart spinifex.target

Check it's up

With Spinifex running, it's worth confirming that Predastore is reachable, and that your client and cluster agree on the region.

Copy Spinifex's CA certificate somewhere your own user can read (and over to any other machine you'll be running clients from). The rest of this guide uses this copy:

sudo install -m 0644 -o "$USER" /etc/spinifex/ca.pem ~/predastore-ca.pem

Point the AWS CLI at the spinifex profile created during setup, which takes care of the credentials, and at your copy of the certificate:

export AWS_PROFILE=spinifex
export AWS_CA_BUNDLE=~/predastore-ca.pem

Then create, list, and remove a test bucket. Listing buckets alone won't catch a region mismatch, which is why the check creates one.

aws s3 mb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 ls s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 rb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1

If all three commands complete without any errors, your Predastore cluster is ready for data. If you see AuthorizationHeaderMalformed ... incorrect region, the region your client is signing for doesn't match the one you passed to spx admin init.

Multi-Node Multi-Drive - Via Spinifex

A caveat before you start: Spinifex currently keeps all of each machine's Predastore data on a single drive. While MinIO is still running, that means picking one of MinIO's drives on each machine, which limits each machine's share of Predastore to that drive's capacity. Predastore's erasure coding still covers a failed machine, and it treats a failed drive as a failed machine, but there's no protection against losing a drive beyond that. If MinIO's drives are already pooled with RAID or ZFS, the pool counts as one drive and keeps its redundancy. Support for spreading Predastore across multiple drives under Spinifex is on its way. Until then, if you need to survive more failed drives than failed machines and can live without IAM, the standalone path is the better fit.

Setting up a multi-node Spinifex cluster involves a fair bit more than Predastore alone, including configuring OVN networking on each server.

Rather than repeating all of it here, I'll point you to Spinifex's multi-node install guide for the full procedure, and only call out the parts that matter for the migration.

Spinifex runs one Predastore blob node per machine, so its erasure coding protects you against losing a machine.

Follow the install guide

Work through steps 1 to 3 of the multi-node guide, which cover installing Spinifex, setting the node IP variables, and configuring OVN networking on each server.

Form the cluster

Step 4 of the guide forms the cluster, with one change; --region needs to match your MinIO clients, on both the init and every join. Spinifex picks the erasure coding for you, using RS(1,1) for two machines and RS(2,1) for three or more.

On server 1, initialize the cluster:

sudo spx admin init --force --node node1 --nodes 3 \
  --bind $SPINIFEX_NODE1 --cluster-bind $SPINIFEX_NODE1 \
  --port 4432 --region us-east-1 --az us-east-1a

While init is still running, join each other server, using the token from init's output. Change --node and the bind addresses to match each server:

sudo spx admin join --force --node node2 \
  --bind $SPINIFEX_NODE2 --cluster-bind $SPINIFEX_NODE2 \
  --host $SPINIFEX_NODE1:4432 --token <token-from-init-output> \
  --region us-east-1 --az us-east-1a

Share one of MinIO's drives

On each machine, Predastore can live on one of the drives MinIO is already using, in a directory of its own alongside MinIO's data, until you're done with MinIO in step 5. Pick the drive with the most free space, as it needs roughly as much room as MinIO is using across all of that machine's drives combined, plus some headroom for re-syncs.

The steps below assume that drive is mounted at /mnt/disk1. On every server, create a directory for Predastore on it. The leading . matters, as MinIO treats any other top-level directory on its drives as a bucket, and will delete it along with that bucket.

sudo install -d -o spinifex-storage -g spinifex -m 0700 /mnt/disk1/.spinifex-predastore

Predastore's service isn't permitted to write anywhere outside /var/lib/spinifex/predastore, so on every server, bind-mount the directory over it. The extra options make sure the service waits for the mount, and won't write to the system disk without it:

sudo chown root:root /var/lib/spinifex/predastore
sudo chmod 0755 /var/lib/spinifex/predastore
echo '/mnt/disk1/.spinifex-predastore /var/lib/spinifex/predastore none bind,nofail,x-systemd.requires-mounts-for=/mnt/disk1,x-systemd.required-by=spinifex-predastore.service,x-systemd.before=spinifex-predastore.service 0 0' | sudo tee -a /etc/fstab
sudo systemctl daemon-reload
sudo mount -a

Start the services

Finish off with steps 5 and 6 of the multi-node guide to start the services and verify the cluster.

As with a single node, Spinifex's certificates only cover each server's own hostname and addresses. If your clients will reach Predastore by any other name, such as s3.example.com, add it on every server, and restart the services to pick it up:

sudo spx admin cert renew --extra-dns s3.example.com
sudo systemctl restart spinifex.target

Check it's up

With the cluster running, it's worth confirming that Predastore is reachable, and that your client and cluster agree on the region. Run the following on server 1.

Copy Spinifex's CA certificate somewhere your own user can read (and over to any other machine you'll be running clients from). The rest of this guide uses this copy:

sudo install -m 0644 -o "$USER" /etc/spinifex/ca.pem ~/predastore-ca.pem

Point the AWS CLI at the spinifex profile created during setup, which takes care of the credentials, and at your copy of the certificate:

export AWS_PROFILE=spinifex
export AWS_CA_BUNDLE=~/predastore-ca.pem

Then create, list, and remove a test bucket. Listing buckets alone won't catch a region mismatch, which is why the check creates one.

aws s3 mb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 ls s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1
aws s3 rb s3://predastore-test --endpoint-url https://s3.example.com:8443 --region us-east-1

If all three commands complete without any errors, your Predastore cluster is ready for data. If you see AuthorizationHeaderMalformed ... incorrect region, the region your client is signing for doesn't match the one you passed to spx admin init.

3. Recreate identities and access

With Predastore up and running, the next step is to give your applications a way in. How you go about this depends on the mode you picked back in step 1.

Standalone

Standalone Predastore has no concept of users or policies. Access is controlled entirely by the [[auth]] entries in predastore.toml, and every one of them is a root credential. So the job here is fairly simple.

Add an [[auth]] entry for each access key your applications currently use against MinIO. You're free to reuse the same key and secret values, which means your applications won't need new credentials when you switch them over.

On a multi-node cluster, make the same change to the config on every machine.

[[auth]]
access_key_id     = "<existing-minio-access-key>"
secret_access_key = "<existing-minio-secret-key>"
account_id        = "100000000001"

Then restart the service (on each machine, for a multi-node cluster) to pick it up:

sudo systemctl restart predastore

Note: If you relied on MinIO policies to restrict what each application could access, there's no standalone equivalent. That's what Spinifex is for.

Via Spinifex

Spinifex implements AWS-compatible IAM, so MinIO's users, groups, and policies map across fairly directly. All of the commands below use the standard AWS CLI, pointed at Spinifex via the profile created during setup.

For more detail on any of them, check out Spinifex's IAM guide.

export AWS_PROFILE=spinifex

Export your MinIO policies

MinIO policies are written in the same JSON format as AWS IAM policies, so in most cases they can be carried over as-is.

List the policies on your MinIO deployment:

mc admin policy list myminio

Then export each of your custom policies to a file:

mc admin policy info myminio <policy-name> --policy-file <policy-name>.json

Before importing them, there are a few things worth checking for, as Spinifex only supports a subset of the IAM policy language:

  • MinIO-specific actions such as admin:* and kms:* mean nothing to Predastore, and should be removed. Spinifex doesn't check action names, so it will accept them without complaint. MinIO's admin:* statements usually have no Resource, though, which Spinifex does reject, with Resource is required.
  • NotAction and NotResource aren't supported yet. Use Action and Resource with an explicit list instead. A statement with only a NotAction or NotResource is rejected with Action is required or Resource is required.
  • Condition keys are limited to aws:SourceIp, aws:SecureTransport, aws:username, aws:userid, aws:PrincipalAccount, aws:PrincipalType, and s3:prefix. The only supported operators are StringEquals, StringLike, IpAddress, and Bool, so conditions using the likes of StringNotEquals will need rewriting.
  • Policy variables are limited to ${aws:username}, ${aws:userid}, and ${aws:PrincipalAccount}. MinIO's ${jwt:*} and ${ldap:*} variables won't carry over.

Create the policies

Create each policy from its exported file:

aws iam create-policy --policy-name <policy-name> --policy-document file://<policy-name>.json

The good news is that Spinifex loudly refuses a policy that uses an unsupported element, condition, or variable. So if one of those slips through, create-policy will fail with a MalformedPolicyDocument error telling you what needs changing. Unrecognised actions are the exception, as mentioned above.

Recreate your users and groups

Each MinIO user and group gets an IAM equivalent, with the same policies attached. You'll need your account ID for the policy ARNs, which this will tell you:

aws sts get-caller-identity

Create each user, and attach their policies:

aws iam create-user --user-name alice
aws iam attach-user-policy --user-name alice \
  --policy-arn arn:aws:iam::<account-id>:policy/<policy-name>

Create each group, attach its policies, and add its members:

aws iam create-group --group-name developers
aws iam attach-group-policy --group-name developers \
  --policy-arn arn:aws:iam::<account-id>:policy/<policy-name>
aws iam add-user-to-group --group-name developers --user-name alice

Issue new access keys

MinIO won't hand over existing secret keys, and Spinifex generates its own, so each user will need a fresh key pair. The secret is only shown once, so make sure to save it somewhere safe.

aws iam create-access-key --user-name alice

Note: Buckets belong to the account that creates them. When you come to copy data across in the next step, use a key belonging to the account you want to own the buckets.

4. Move the data

Now for the main event. To copy your data from MinIO to Predastore, I recommend using rclone.

You might be tempted to reach for mc mirror, and it does work, but with MinIO archived, mc is no longer maintained and its official downloads have been taken down. rclone on the other hand is actively maintained, packaged for pretty much every distro, and lets you set the region explicitly on both ends.

Install rclone

Install it on a machine that can reach both MinIO and Predastore:

sudo -v ; curl https://rclone.org/install.sh | sudo bash

Configure both ends

rclone needs a remote for each side. Add the following to ~/.config/rclone/rclone.conf, filling in your own endpoints and keys. force_path_style is required, as Predastore only supports path-style requests, and region needs to match the one you set up Predastore with.

If you went with Spinifex, use a key belonging to the account you want to own the buckets. The spinifex profile's own key, in ~/.aws/credentials, belongs to the account created during setup.

[minio]
type = s3
provider = Minio
endpoint = https://minio.example.com:9000
access_key_id = <minio-access-key>
secret_access_key = <minio-secret-key>
region = us-east-1

[predastore]
type = s3
provider = Other
endpoint = https://s3.example.com:8443
access_key_id = <predastore-access-key>
secret_access_key = <predastore-secret-key>
region = us-east-1
force_path_style = true

Confirm both remotes work by listing their buckets:

rclone lsd minio:
rclone lsd predastore: --ca-cert ~/predastore-ca.pem

Copy each bucket

Copy a bucket by creating it on the Predastore side, then filling it with rclone sync. rclone only creates a bucket when it has something to put in it, so without the mkdir an empty bucket would quietly get left behind.

--checksum compares objects by their hash rather than their modification time, and --metadata carries across user metadata (x-amz-meta-*) and Content-Type.

rclone's --ca-cert flag applies to both remotes at once, not just Predastore. If your MinIO is served over HTTPS, passing Predastore's certificate on its own will stop rclone from trusting MinIO, so pass the system's CA bundle alongside it (on RHEL and Fedora, that's /etc/pki/tls/certs/ca-bundle.crt):

rclone mkdir predastore:<bucket> --ca-cert ~/predastore-ca.pem
rclone sync minio:<bucket> predastore:<bucket> \
  --checksum --metadata --progress \
  --ca-cert /etc/ssl/certs/ca-certificates.crt --ca-cert ~/predastore-ca.pem

Note: rclone sync makes the destination match the source, which means deleting anything in the destination that isn't in the source. Always name the bucket on both sides, and never run it against the root of the remotes, as that will clear out anything in Predastore that MinIO doesn't know about. If in doubt, add --dry-run first to see what it would do.

To copy every bucket in one go:

for bucket in $(rclone lsf --dirs-only minio: | tr -d /); do
  rclone mkdir "predastore:$bucket" --ca-cert ~/predastore-ca.pem
  rclone sync "minio:$bucket" "predastore:$bucket" \
    --checksum --metadata --progress \
    --ca-cert /etc/ssl/certs/ca-certificates.crt --ca-cert ~/predastore-ca.pem
done

Check the copy

Check each bucket with rclone check. --download reads every object from both sides and compares the actual contents. It's slow for large buckets, but it's the most reliable way to be sure.

rclone check minio:<bucket> predastore:<bucket> --download \
  --ca-cert /etc/ssl/certs/ca-certificates.crt --ca-cert ~/predastore-ca.pem

A successful check reports 0 differences found.

5. Switch clients over and check the result

With the data across, all that's left is to point your applications at Predastore. It's worth doing this in a sensible order, so that nothing written to MinIO in the meantime gets left behind.

Stop writes to MinIO, and run a final sync

Stop your applications from writing to MinIO (or put them into maintenance mode), then run the copy from step 4 one last time to pick up anything that changed since.

Compare the two sides

Before switching anything over, check that both sides agree on the number of objects and the total size of each bucket:

rclone size minio:<bucket>
rclone size predastore:<bucket> --ca-cert ~/predastore-ca.pem

Update your clients

For most clients, switching over comes down to a handful of settings:

  • Endpoint - https://<predastore-host>:8443. Predastore only speaks HTTPS, so any http:// endpoints will need updating along with the host and port.
  • Region - Whatever you set during setup in step 2. A mismatch shows up as an AuthorizationHeaderMalformed error on every bucket and object request.
  • Credentials - Unchanged if you reused your MinIO keys on a standalone deployment. Otherwise, the new keys from step 3.
  • Addressing style - Predastore only supports path-style requests (https://host:8443/bucket/key), not virtual-hosted ones (https://bucket.host:8443/key). Most SDKs have an option for this, such as addressing_style in boto3, forcePathStyle in the JavaScript SDK, and UsePathStyle in the Go SDK.
  • Certificate trust - If you're using a self-signed certificate, clients need to trust it. That's the ~/predastore-ca.pem copy made when you checked the cluster was up in step 2.

For example, an AWS CLI profile for Predastore in ~/.aws/config looks something like this:

[profile predastore]
region = us-east-1
endpoint_url = https://s3.example.com:8443
ca_bundle = /home/<you>/predastore-ca.pem
aws_access_key_id = <access-key>
aws_secret_access_key = <secret-key>
s3 =
  addressing_style = path

Note: Any presigned URLs generated before the switch point at MinIO, and won't carry over. Your applications will need to generate new ones against Predastore.

Keep MinIO around for a little while

Once your applications are happily talking to Predastore, I'd recommend leaving MinIO running in a read-only state for a while before decommissioning it. That way, if anything turns out to be missing, the original is still there to fall back on.

Retire MinIO

Once you're confident nothing is missing, stop MinIO for good, and delete its data to hand the space over to Predastore. Predastore carries on running throughout, and there's nothing to restart. Under Spinifex, only the drive holding Predastore's data gains anything for now, and the rest are free for other uses.

If MinIO was pointed at a directory on each drive, such as /mnt/disk1/minio, delete that directory on every drive:

sudo rm -r /mnt/disk1/minio

If MinIO was pointed at the drive itself, delete its .minio.sys directory, along with a directory for each of your buckets, naming them explicitly:

cd /mnt/disk1 && sudo rm -r .minio.sys <bucket> <bucket> ...

Note: Resist the temptation to reach for rm -r /mnt/disk1/*. It misses .minio.sys, and if Predastore's directory is ever missing its leading ., it'll take your freshly migrated data with it.

Closing thoughts

Hurray! Hopefully that went smoothly for you. If it didn't, raise a GitHub issue so we can help you out.

We've got plenty of features and improvements planned for Predastore, so be sure to keep your eyes on the release page.

I've got other articles in the works as well, focusing more on the design rationale behind the system. Stay tuned for those. That's all for this one!

#Predastore#MinIO#S3#Object Storage#Migration