How to set up replicators¶
Replicators sync instances across LXD cluster links. You can prepare for active-passive disaster recovery by setting up replicators to periodically copy instances from a primary cluster to a secondary cluster. The primary (active) cluster handles all workloads while the secondary (passive) cluster remains ready to take over if the primary cluster fails.
LXD supports this strategy using project replicators over a cluster link.
Prerequisites¶
Before setting up replicators:
Two LXD clusters must be initialized (the primary and secondary clusters).
You need sufficient permissions on both clusters to create cluster links and manage projects.
A bidirectional cluster link must be established between the two clusters.
Network connectivity must exist between the clusters.
Additional prerequisites for Ceph RBD mirroring¶
By default, a replicator copies the data of every instance over the cluster link. If both clusters keep the project on Ceph RBD storage pools, you can set up the replicator to use Ceph RBD mirroring to copy the data. LXD then sends only the instance and volume records over the cluster link. See Replicators with Ceph RBD mirroring for details.
In this case, your setup must also meet these prerequisites:
Both clusters support the storage_ceph_replicator API extension.
Each cluster uses its own Ceph cluster.
The LXD storage pool has the same name on both clusters, and so does the OSD pool behind it. Ceph mirrors a volume to the OSD pool of the same name, and the instance records refer to the storage pool by name.
The secondary cluster’s storage pool is a regular storage pool. Do not create it with
source.recover, which is meant for storage replication of a whole OSD pool.RBD mirroring is enabled in
imagemode on the OSD pool in both Ceph clusters. Inpoolmode, Ceph would mirror every volume in the OSD pool, including the volumes of other projects.The two Ceph clusters are peers of each other in both directions (
rx-tx), and therbd-mirrordaemon runs in both. After a failover, the clusters swap roles, so each Ceph cluster must be able to receive data from the other.The profiles that the project’s instances use exist on the secondary cluster with the same devices, including a root disk device on the mirrored storage pool.
Visit the Ceph documentation on RBD mirroring for details about how to enable mirroring and add the peers.
Create projects for replication¶
Set up projects with the same name on both clusters, and set the replica.cluster configuration key on both projects to identify the cluster links allowed to replicate instances between them.
Note
At setup, the replica.cluster configuration key is only required for the standby project on the secondary cluster. During failover and failback, however, the key is required to promote the project on the primary cluster from standby to leader mode.
On the primary cluster, create a project and point it at the secondary cluster:
lxc project create <project_name> --config replica.cluster=<secondary_cluster_link_name>
Expand the Project drop-down and select + Create project at the bottom.
Enter a name and optionally a description for the new project.
Go to the new project’s configuration and select the Replication tab.
Under Replica cluster, select the cluster link established between the primary and secondary clusters.
On the secondary cluster, create a project with the same name and configure it to accept replication from the primary cluster:
lxc project create <project_name> --config replica.cluster=<primary_cluster_link_name>
Expand the Project drop-down and select + Create project at the bottom.
Enter the same name as in the previous step, and optionally a description for the new project.
Go to the new project’s configuration and select the Replication tab.
Under Replica cluster, select the cluster link established between the primary and secondary clusters.
Configure storage pools for Ceph RBD mirroring¶
Skip this section if you do not use Ceph RBD mirroring to copy the data.
On each cluster, the ceph.replicator.<project> key on the storage pool names the peer Ceph site, which is the Ceph cluster that the other cluster uses.
A project with this key set on one of its storage pools is a mirrored project.
To see the site names, run the following command against either Ceph cluster:
rbd mirror pool info <osd_pool_name>
On the primary cluster, set the key to the site name of the secondary cluster’s Ceph cluster, and set
ceph.rbd.clone_copytofalse:lxc storage set <pool_name> ceph.rbd.clone_copy=false ceph.replicator.<project_name>=<secondary_ceph_site_name>
On the secondary cluster, set the key to the site name of the primary cluster’s Ceph cluster, and set
ceph.rbd.clone_copytofalse:lxc storage set <pool_name> ceph.rbd.clone_copy=false ceph.replicator.<project_name>=<primary_ceph_site_name>
Important
Set ceph.rbd.clone_copy to false before you create any container on the storage pool.
Ceph cannot mirror a volume that is a clone of its image.
This setting applies to all projects that use the storage pool.
The project must exist before you can set ceph.replicator.<project>.
While the key is set, the following rules apply to the project:
Every instance, and every custom volume attached only to one instance, must be on a storage pool that has the key. Otherwise, the replicator refuses to run.
Ceph mirrors all instance volumes and custom volumes that the project holds on the storage pool. The secondary cluster, however, only receives records for instances and for custom volumes attached only to one instance. Do not attach a custom volume to more than one instance in a mirrored project: unlike with other replicators, you cannot create it in advance on the secondary cluster.
You cannot rename the project.
To delete the project, unset the key first, or delete the project with
--force, which also removes the key.
Prepare authentication¶
Replicators communicate over cluster links, so the linked cluster identities must be granted the permissions they need on the replicated project. Configure these permissions using authentication groups and Manage permissions.
For project replication, the cluster-link identity on each cluster typically needs at least these permissions:
operatoron the replicated project, so the cluster link can perform instance replicationcan_editon the replicated project, so replica project configuration can be validated and updated as part of the workflow
For example, if the replicated project is called myproject, you can prepare an authentication group named replicators on each cluster:
lxc auth group create replicators
lxc auth group permission add replicators project myproject operator
lxc auth group permission add replicators project myproject can_edit
Then, on each cluster, add the cluster links to that authentication group, as described in Manage cluster link permissions. (You can also assign cluster link identities to an authentication group when you create the links, as detailed in How to create cluster links.)
Create a replicator¶
On the primary cluster, create a replicator that targets the secondary cluster. The replicator must exist before the project on the primary cluster can be promoted, because the replicator identifies the source of replication. Each cluster link can be targeted by at most one replicator per project; creating or updating a replicator to target a cluster link used by another replicator in the same project fails with a conflict error.
lxc replicator create <replicator_name> cluster=<secondary_cluster_link_name> --project <project_name>
For example:
lxc replicator create my-replicator cluster=lxd-standby --project myproject
You can also create a replicator with a schedule:
lxc replicator create my-replicator cluster=lxd-standby schedule="@daily" --project myproject
Click Clustering in the navigation sidebar, then select Replicators from the expanded drop-down list.
Click on the + Create replicator button to open the side panel.
Select the cluster link established between the primary and secondary clusters. Enter a name and optionally a description for the new replicator. Select the project that you configured in the previous steps. You can also enter a schedule.
Click Create.
See Replicator configuration for all available configuration options.
Configure project replica modes¶
On the secondary cluster, demote the project to
standbymode. Instances cannot be created or started in a standby project. The project must be promoted toleaderduring a failover before instances can be started.lxc project demote-replica <project_name>
Under Replica mode, click Demote to standby.
On the primary cluster, promote the project to
leadermode:lxc project promote-replica <project_name>
Go to the leader project’s configuration and select the Replication tab.
Under Replica mode, click Promote to leader.
Promotion from an unset replica mode to leader mode requires the project to have at least one replicator.
LXD also validates that all target projects on clusters referenced by the project’s replicators are in standby mode before allowing promotion.
This ensures that new instances are not created on a standby project between replicator runs.
If a target cluster is unreachable, promotion still proceeds to allow disaster recovery scenarios where the target may be offline.
Force-promote a project to skip these checks, for example when migrating a deployment that was set up before replicators were required for promotion.
Run a replicator¶
To manually trigger a replicator run:
Use the following command on the primary cluster:
lxc replicator run <replicator_name>
On the primary cluster, click Clustering in the navigation sidebar, then select Replicators from the expanded drop-down list.
Click on the “run” button at the end of the replicator’s row.
Alternatively, click on a replicator name to view its detail page, then click on the Run button in the header.
This syncs all instances in the source project to the secondary cluster, along with custom volumes that are attached only to those instances.
If the project is mirrored with Ceph RBD, the run completes only after the secondary cluster’s Ceph cluster has received the data. The first run copies every volume in full, so it can take a long time. If Ceph needs more than one hour, the run fails, but Ceph continues to copy the data. In this case, run the replicator again later.
To schedule replication automatically, set the schedule configuration key with a cron expression:
lxc replicator set <replicator_name> schedule="0 0 * * *"
In the Create replicator or Edit replicator side panel, enter a cron expression in the Schedule input box.
For a mirrored project, set the schedule only after the first run has succeeded. A scheduled run that starts while Ceph is still copying the volumes in full takes a new mirror snapshot of every volume, and can fail in the same way.
Snapshot before replication¶
Each replicator run performs an incremental instance sync to the secondary cluster using
the equivalent of lxc copy --refresh. This transfers only the data that has changed since the last sync,
using any existing snapshots as a reference point to minimize the amount of data transferred.
Before the incremental copy, LXD creates a point-in-time snapshot of each source instance. This gives the copy operation a consistent reference point, which reduces the amount of data transferred on each sync and provides a rollback point on the source in case anything goes wrong during replication.
Snapshot naming and expiry are controlled entirely by the instance’s own configuration (for
example snapshots.pattern and
snapshots.expiry), or by the profile applied to the
instance. The replicator does not impose its own naming scheme.
If an instance already has a snapshots.schedule set at
the instance or profile level, the replicator skips creating a new snapshot and reuses the
most recent existing snapshot as the reference point for the incremental copy instead.
Note
Snapshots created by replication accumulate over time. Use snapshots.expiry on the instance or
profile to automatically prune them, or delete them manually with lxc snapshot delete.
Next steps¶
Once replicators are running, see How to manage replicators to view, configure, or delete replicators, and How to perform disaster recovery with replicators to fail over to the secondary cluster if the primary cluster becomes unavailable.