Storage

From canasta

This page describes how Canasta handles persistent storage: what state it keeps, how each orchestrator (Docker Compose, Kubernetes) provisions that state, when shared (RWM-capable) storage is required, and which storage backends Canasta supports out of the box.

What Canasta stores

A running Canasta instance has six categories of persistent state:

State Lives on
Wiki content directories โ€” extensions, skins, images, public_assets PVCs (K8s) or instance-directory bind mounts (Compose)
Wiki database (when bundled) A db-data PVC managed by a StatefulSet (K8s), or the mysql-data-volume Docker volume (Compose). Skipped entirely when USE_EXTERNAL_DB=true โ€” see Help:External database.
Search index (when Elasticsearch enabled) An es-data PVC managed by a StatefulSet (K8s), or the elasticsearch Docker volume (Compose).
TLS state (Compose only) The caddy-data Docker volume holds Caddy's Let's Encrypt account and issued certificates. K8s instances delegate TLS to cert-manager + ingress; Caddy runs with no persistent state and has no equivalent PVC.
Logs (when observability stack enabled) Shared volumes between web/db/caddy and the log-shipping containers. See Help:Observability for the topology.
Instance config โ€” .env, values.yaml, per-wiki settings The instance directory at ~/canasta/<id>/ on the operator workstation (or wherever --path points). Not on the cluster.

The instance config is on the operator host, not the cluster. If the operator host fails, the cluster keeps running but you lose the ability to issue further canasta commands until the directory is restored. See Help:GitOps for keeping it in version control, or back it up alongside the wiki content.

Storage on Compose

Compose instances use two storage mechanisms:

  • Bind mounts from the instance directory: ./extensions, ./skins, ./images, ./public_assets, ./config are mounted into the web container at the equivalent /var/www/mediawiki/ paths. Editing files in the instance directory directly modifies what the running container sees.
  • Named Docker volumes for the bundled services: mysql-data-volume, caddy-data, elasticsearch, plus log-shipping volumes when the observability profile is active. These are managed by Docker Engine โ€” typically under /var/lib/docker/volumes/ on Linux hosts.

Compose has no concept of a StorageClass; whatever the host filesystem offers is what you get. Backups (see Help:Backup and restore) capture both the bind-mounted directories and the named volume contents into a single Restic snapshot.

Disk space and relocating container storage

Container images and named volumes share whatever filesystem the container runtime's storage root lives on โ€” /var/lib/docker for rootful Docker, ~/.local/share/containers/storage for rootless Podman. On hosts where that path sits on a small root partition (a common VPS/cloud default is a ~50 GB root volume with data on a separate disk), it fills up quickly: the Canasta image alone is ~2.3 GiB, and the database, search index, and Caddy volumes grow from there. A full root partition destabilizes the whole host, not just the wiki.

canasta create runs a best-effort pre-flight check and prints a warning when the container-storage filesystem has less than 20 GiB free. The warning is advisory โ€” it does not stop the install. To put container storage on a larger volume, configure the runtime before the first canasta create:

  • Docker: set "data-root": "/path/on/large/volume" in /etc/docker/daemon.json, then restart Docker (sudo systemctl restart docker).
  • Rootless Podman: set graphroot = "/path/on/large/volume" under [storage] in ~/.config/containers/storage.conf.

Relocating after images and volumes already exist means migrating that data (or recreating the instance), so it is far easier to set this up first.

Storage on Kubernetes

Kubernetes instances use PersistentVolumeClaims for everything stateful. The cluster's CSI drivers and StorageClasses control the actual provisioning.

Which StorageClass gets used

Canasta picks the StorageClass for a new instance in this order, first match wins:

  1. --storage-class <name> on canasta create.
  2. The defaultStorageClass setting in the controller's registry (~/Library/Application Support/canasta/conf.json on macOS, ~/.config/canasta/conf.json elsewhere). Set this once per controller and skip the flag.
  3. The cluster's default StorageClass (whichever one is annotated storageclass.kubernetes.io/is-default-class: "true").

The registry value is sticky across instances. If the same controller has been used against multiple clusters with different defaults (e.g., k3s's local-path and EKS's gp2), the registry locks in whichever was set last; it does not re-read the cluster's default per-instance. Pass --storage-class explicitly when switching clusters, or update the registry once with the new value.

To inspect what's available on the current cluster:

canasta storage list

ReadWriteOnce vs ReadWriteMany

The four content PVCs default to ReadWriteOnce (RWO). RWO PVCs can mount on only one node at a time, which is fine for single-replica web on a single-node cluster, or for multi-replica web where Kubernetes can schedule all replicas onto the same node.

For multi-node multi-replica web โ€” where pods on different nodes need to read and write the same content โ€” pass --access-mode ReadWriteMany at canasta create time. RWM declares the PVCs RWM-capable and matches them to a StorageClass whose CSI driver supports concurrent multi-node mounts. Common RWM backends:

  • Network filesystems exported from a shared host or appliance (NFS, SMB, CephFS).
  • Managed RWM services from cloud providers (e.g., AWS EFS, Azure Files, Google Cloud Filestore).
  • Vendor CSI drivers that explicitly advertise RWM (NetApp Trident, Portworx, Longhorn-with-RWM-enabled).

The single-node defaults โ€” k3s's local-path, EKS's gp2 โ€” are RWO-only. Asking for RWM access mode against an RWO-only StorageClass produces PVCs stuck in Pending. See Help:Multi-node Kubernetes for the topology context and examples.

Turnkey storage helpers

For two common shared-storage backends, Canasta installs the CSI driver and registers the StorageClass in one command:

Backend Command
NFS (server on any reachable host, including a Canasta-managed node) canasta storage setup nfs --host <node> --share /srv/nfs/canasta
AWS EFS (managed NFS, pre-provisioned filesystem) canasta storage setup efs --host <node> --filesystem-id fs-...

Both commands target a --host that has a working kubectl against the cluster โ€” typically the K8s control-plane node. canasta storage setup nfs --install-server additionally installs the NFS server package on that host (nfs-kernel-server on Debian/Ubuntu, nfs-utils on RHEL/Fedora), creates the share directory, and exports it to the cluster's nodes โ€” see NFS server exports below. With --server <address> instead, Canasta uses an NFS server you already run: it checks that the server answers on port 2049 but does not change its exports, so make sure they allow the cluster's nodes to mount the share.

NFS server exports

ℹ️ Note: This section describes Canasta CLI 4.21.0 and later. Earlier releases exported the share to any host.

With --install-server, the share is exported only to the cluster's node addresses (each node's internal and external IP) and pod networks, never to any host. Canasta writes these exports to /etc/exports.d/canasta.exports and leaves /etc/exports to you, apart from removing an any-host line for the same share. The export options are rw,sync,no_subtree_check,no_root_squash.

The client list is kept current by a systemd timer, canasta-nfs-exports-sync.timer, which re-reads the cluster's nodes every minute, so nodes that join or leave the cluster are picked up without re-running canasta storage setup nfs. The sync runs as root and reads the cluster through the kubeconfig that was in effect for the SSH user when you ran the setup ($KUBECONFIG, else ~/.kube/config), falling back to /etc/rancher/k3s/k3s.yaml.

If the sync cannot list the cluster's nodes โ€” for example while the Kubernetes API is unreachable โ€” it leaves the existing exports unchanged rather than withdrawing them. When that happens during canasta storage setup nfs and the share ends up not exported, the setup fails. To see why the sync is failing, run it by hand on the NFS host:

sudo /usr/local/sbin/canasta-nfs-exports-sync

Because the share is exported with no_root_squash, keep port 2049 closed to the internet as well; the client list limits who can mount it, not who can reach the port.

Shares set up by earlier releases. canasta upgrade checks the host of each Kubernetes instance for a share that an earlier canasta storage setup nfs --install-server exported to any host, and moves it to the cluster-only export. If the sync cannot list the cluster's nodes at that point, the upgrade keeps the old export, prints a warning, and carries on; fix the cause and re-run canasta upgrade.

Uninstalling Kubernetes. canasta uninstall k8s on a k3s node that also serves such a share withdraws the export once the kubeconfig the sync reads no longer exists, and removes the sync's timer, service, and script. The share's contents are left on disk. If the sync follows a kubeconfig that is still present, the exports are left in place.

For other RWM backends (CephFS, Portworx, vendor CSI drivers), follow your CSI driver's standard install instructions and register the StorageClass via kubectl apply -f; Canasta only needs the StorageClass to exist when canasta create --storage-class <name> runs.

Object storage

Two distinct object-storage roles are worth distinguishing:

  • Backups use object storage as the Restic repository โ€” the actual destination of snapshot data. This is fully supported via the RESTIC_REPOSITORY environment variable. See Help:Backup and restore for the supported backends (S3, Azure Blob, Backblaze B2, GCS, OpenStack Swift, etc.) and configuration.
  • Wiki uploads as object storage โ€” having MediaWiki write user-uploaded files directly to S3 instead of the images PVC โ€” is not currently supported by the bundled chart. Uploads go to the images PVC and are reachable via the wiki's normal URL paths. Operators who want object-storage-backed uploads can configure MediaWiki's $wgFileBackends manually in their settings, but this is operator-managed configuration outside Canasta's supported envelope.

Sizing

The default per-PVC size in the chart (1 GiB for each of extensions, skins, images, public_assets) is sized for a small bootstrap. Real wikis with media uploads, third-party extensions, or large skins will need to grow these โ€” typically images grows fastest. Edit the per-instance values.yaml on the target host (persistence.images.size, etc.) and run canasta restart to apply. Some CSI drivers (most cloud-managed ones, and NFS) support online expansion; check your provider's docs.

See also