Skip to main content

Why Cloud-Assisted Recovery Is Becoming More Attractive to Mainframe Shops

Cloud infrastructure experts explain how services such as AWS CloudFormation and S3 can help enterprises build scalable resilience layers without moving core IBM Z workloads

TechChannel Data Management

Implementing cloud services for operational resilience does not necessarily mean moving core mainframe workloads off IBM Z. In regulated industries such as banking, insurance and government, on-premises servers often remain the systems of record, but enterprises can use cloud services to strengthen disaster recovery.

Given tighter regulatory expectations for geographic recovery, the cost of maintaining secondary data centers and the practical reality that many business-critical workloads are unlikely to be migrated in the near term, a cloud-assisted approach can help to make mainframe recovery faster and more cost efficient.

In this model, AWS provides data replication recovery orchestration and scalable recovery infrastructure while critical production workloads continue running on local servers.

“Using the cloud as a resilience layer can give organizations another path to improve mainframe disaster recovery without requiring a full migration,” Scott Petry, principal, Cloud Engineering, Data & Analytics at PwC US, tells TechChannel. “Instead of maintaining dedicated secondary infrastructure that may sit largely idle, organizations can use more on-demand cloud capacity and potentially test recovery scenarios more frequently.”

Cloud-based disaster recovery changes the economics of backup systems by focusing on data management.

“What’s interesting about disaster recovery in the public cloud is that’s all you need to do, sending the data, because if you think about doing disaster recovery in the on-prem world, having the data itself is not going to help—you need some compute and networking and everything else to restart your systems,” Fred Lherault, Field CTO EMEA and Emerging Markets at Everpure, tells TechChannel.

Automating Recovery Into AWS

Enterprises can replicate critical data to the cloud and provision the recovery environment only when a failover is required. Infrastructure-as-code and automation tools, such as the AWS Cloud Development Kit (CDK), as well as Terraform or Ansible, can provision compute and networking resources automatically on demand through AWS CloudFormation.

“If you’ve done a good job of automating everything, you only need to pay for the data … to be replicated and stored in an availability zone, you’re not paying for compute, you’re not paying for any other services to deploy in that region until you actually need it,” Lherault explains.

Enterprise teams can build and manage this orchestration internally or use managed recovery services. Lherault describes a scenario is which customers running VMware on-premises replicate their data into AWS and, when they need to perform a disaster recovery test or execute a real failover, the service provider can automatically provision all the instances they need.

The primary cost is data replication and storage, not a fully provisioned recovery site, shifting the emphasis from infrastructure ownership to recovery readiness.

“That’s one of the clever ways of using the cloud and that is enabled by the elasticity you get from the public cloud,” Lherault says.

AWS Elastic Disaster Recovery is increasingly being used as a resilience layer for mainframe environments rather than as a migration mechanism. The service continuously replicates data from on-premises systems into AWS, helping teams to achieve defined recovery point objectives while keeping production workloads running in mainframe systems.

Elastic Disaster Recovery provides an automated framework for connected systems and data, including API gateways, integration platforms, analytics services and distributed applications that depend on z/OS data. It can support broad hybrid disaster recovery scenarios in which only selected workloads fail over to AWS while core mainframe processing remains on-premises.

“The approach can also provide geographic flexibility and help organizations build the connectivity, automation and operational experience needed for future modernization efforts,” Petry notes.

Amazon Simple Storage Service (S3) plays a role in cloud-assisted resilience, providing a durable layer for backup and recovery data. Unlike traditional block storage, Amazon S3 is an object storage service designed for long-term, large-scale retention and geographically distributed data protection.

S3 can store backups, snapshots, archives and artifacts for disaster recovery, while features such as immutable storage options and retention controls help protect data against accidental deletion or ransomware.

The Networking and Replication Equation

Cross-region replication can further increase resilience by maintaining copies of critical recovery data in a geographically separate AWS Region.

For many regulated organizations, resilience planning must account for regional disruptions. Financial services regulations such as the European Union’s Digital Operational Resilience Act (DORA) increasingly require firms to demonstrate that critical services can continue operating even if an entire geographic area is affected by a disaster.

Meaningful separation, often across hundreds of miles, between primary and recovery environments typically means replicating recovery data to a different AWS Region rather than relying on multiple Availability Zones (AZs) within the same region.

Replicating data across AZs or Regions improves resilience, but it also increases data transfer requirements and costs.

“Organizations may underestimate the cost and operational complexity involved,” Petry says. “Continuously replicating large volumes of mainframe data can introduce storage, network and data-transfer costs that should be modeled upfront.

The challenge is not only financial. Recovery environments must also be capable of supporting the operational needs of the workloads they protect.

“The trade-off is that cloud infrastructure should not be assumed to provide mainframe-class performance automatically,” Petry says. “Organizations must validate whether the recovery environment can support the required transaction volumes, latency and operational characteristics, and whether it can meet established recovery time and recovery point objectives.”

In mainframe-connected applications with tightly controlled throughput and response times, a recovery environment that can restore data but cannot support critical workflows may not meet the organization’s resilience needs.

“For many organizations, the cloud can be an effective resilience layer, but only when the recovery design is tested against the actual requirements of the workloads rather than treated as a direct replacement for the mainframe environment,” Petry says.

From Data Replication to Operational Readiness

A successful recovery strategy requires organizations to validate that data, applications and operational processes will function correctly when they are needed.

“Data replication, automated provisioning and infrastructure as code are foundational, but they are not enough on their own,” Petry notes. “Organizations also need to prove that the mainframe workloads can be executed, validated and returned safely to production.”

The first challenge is validating that the recovery environment can run the workloads it is designed to support. “Replicating data to the cloud provides limited value unless the target environment can run the applications and transactions that depend on it,” Petry explains. “Organizations need a proven approach for COBOL or PL/I applications, CICS or IMS transactions and Db2 workloads, whether through emulation, recompilation or a managed service.”

Recovery environments must also meet the same requirements for security, compliance, audit, data sovereignty and licensing as production systems.

Data fidelity is another key prerequisite. “Mainframe data may include EBCDIC encoding, packed-decimal fields, VSAM files, and IMS structures that do not always translate cleanly to modern platforms,” Petry says. “Automated validation and reconciliation should confirm that the data in the recovery environment matches the system of record accurately and completely.”

Finally, teams need to plan for the return to normal operations with a defined and tested failback process. “The most common misconception is that successful data replication means the organization is ready to recover in the cloud. Replication is not the same as recoverability,” Petry warns. “Returning workloads and reconciled data to the production mainframe can be more complex than the initial failover.”

“A disaster recovery plan is proven only when the organization can execute the workloads, meet its recovery objectives, and complete both failover and failback under realistic conditions.”

Mainframe Resilience Without Migration

Testing these processes regularly, before an outage occurs, turns cloud disaster recovery from a theoretical capability into a dependable resilience strategy.

“The organizations that get this right treat cloud-based disaster recovery as an operational capability that must be engineered, exercised and continually improved, not as a box that is checked once replication is enabled,” Petry concludes.

The future of mainframe resilience is increasingly separate from migration. Enterprises can take advantage of cloud capabilities without moving core IBM Z workloads, using AWS as a scalable resilience layer that strengthens recovery, automation and geographic protection.


Key Enterprises LLC is committed to ensuring digital accessibility for techchannel.com for people with disabilities. We are continually improving the user experience for everyone, and applying the relevant accessibility standards.