BCDR in digital asset custody: why a standby environment is needed and how to prepare one
The reliability of custody infrastructure is defined not only by how it performs in normal conditions, but also by what happens when the usual order is disrupted. If the primary platform becomes unavailable, how do you continue operations, preserve the necessary checks and approvals, and meet your obligations to clients? It is worth having the answer prepared before a disruption occurs.
This is the task BCDR addresses: business continuity and disaster recovery. While many providers supply a disaster recovery kit to recover key material or data, business continuity, ensuring that operations can proceed safely without interruption during a failure, remains an essential and distinct requirement. In digital asset custody, true operational resilience requires preparing not only for the restoration of systems, but for the capability to transact safely while primary operations are disrupted.
The Vault now offers a standby custody environment for institutions that already use another platform as their primary. It allows the alternative infrastructure to be provisioned in advance, the rules for executing operations to be configured, and the team's actions to be rehearsed against the scenario in which the primary service is unavailable.
Why a backup of the keys is not enough
The ability to execute a transaction depends on more than key material. It also requires authorised operators, approvals, verification of the recipient address and observance of limits. All of this is delivered by a combination of technology and the institution's own internal procedures.
If signing and approval depend on a platform that is unavailable, holding a key share or a backup of key material does not in itself guarantee that operations continue. What matters is knowing where and how signing will be restored, and which rules will apply in the standby environment.
BCDR therefore covers infrastructure, people and processes together. Its purpose is to establish which operations have to be restored first, within what time, and who is responsible for each stage.
What The Vault offers
The standby environment is prepared in advance and stays outside day-to-day operations until the institution activates it under an agreed recovery scenario.
Segregated wallets are created inside it using threshold MPC, a distributed signing technology. Operator roles and approval quorums are configured: who may initiate a transaction, and how many confirmations are required for it to be executed. Limits, allowlisted addresses and screening rules are applied before signing.
This is working infrastructure that the team becomes familiar with before an incident. It is there to reduce the volume of technical and organisational steps that would otherwise have to be carried out in the middle of a disruption.
It is important here to distinguish the readiness of a standby platform from the availability of assets. Creating a second environment does not in itself move funds out of the primary one. The recovery plan has to set out separately how the institution will reach its existing assets if the primary provider is unavailable. That scenario, alongside the operation of the standby environment, needs to be tested in advance.
How preparation works
The first step is to define the scenarios the institution wants to be ready for. A temporary loss of service and the need to change provider entirely call for different actions. For each scenario it is important to set the tolerable interruption, the priority operations and the conditions for activating the standby environment.
The next step: configuring the infrastructure and the integration. The Vault provides a Fireblocks-compatible API, which allows a standby connection to be prepared for institutions using that platform. The compatibility of specific operations and any changes required are verified during implementation.
In parallel, responsibilities are assigned: who takes the decision to switch over, who executes and approves operations, who verifies the result. These arrangements directly align with broader institutional custody trends. According to findings in the 2026 Institutional Investor Digital Assets Survey, prepared by Coinbase and EY-Parthenon, 61% of investors currently employ a multi-custodian framework, typically using two to three custodians or a mix of outsourced and internal custody, primarily to mitigate single-point-of-failure risks. Implementing a standby environment offers the operational continuity of a multi-custodian setup while preserving unified risk and governance standards.
The standby environment has to support the agreed way of working, and it should not become an exception to the security rules.
Why rehearsals matter
A plan becomes useful in practice once the team has established that it can execute it.
Failover rehearsals are run with the named operators. They make it possible to check access, approvals and the execution of the operations covered by the scenario, and to measure the actual recovery time. The results are documented, and the issues found are used to refine the plan and the configuration.
This kind of check surfaces obstacles that are hard to see in a document: an approver who cannot be reached, missing permissions, or discrepancies between the rules of the primary and the standby environments.
In a real incident the team follows a procedure it has already proved. The aim is to restore the necessary operations while retaining control over who performs them and on what terms.
Why this matters now
The operational resilience of custody is under European supervisory focus. In July 2026 ESMA announced a Common Supervisory Action on CASPs covering key and storage management, transaction controls, incident detection and response, and dependencies on third-party providers. The exercise runs from the second half of 2026 through the first half of 2027. More in the ESMA announcement.
The practical task exists independently of the supervisory timetable. An institution needs to understand which capabilities remain available to it if its primary provider becomes unavailable, and what evidences its team's readiness to carry on operating.