Introducing Pull For Advanced Data Replication

author-image
DQChannels Bureau
New Update

The challenge in today's business for data protection is to reduce risk and
increase business resilience, while also reducing costs and increasing
efficiency. Coupled with the disk-based journaling strategy, use of a pull-based
replication process can pull the data from the primary storage system's
journal volume across the link and write it to the journal volume at the
receiving site

Advertisment

While trying to address business continuity needs for today, enterprises must
respond to new business drivers such as round-the-clock operations, higher
service-level expectations, closer regulatory scrutiny with the emergence of
stringent out-of-region data protection requirements, and increased sensitivity
to loss of data and information assets. Hence the challenge is to reduce risk
and increase business resilience, while also reducing costs and increasing
efficiency.

In many industries and geographies, government regulations require companies
to have effective business continuity plans that will enable them to protect
information assets and maintain their service capabilities in spite of local or
regional disaster.

Factors
that are crucial to charting out business continuity plans include data
replication with guaranteed integrity and consistency, scope of definition of
data that needs replication, better RTO (Recovery Time Objective) and RPO
(Recovery Point Objective). In order to meet this, organizations follow many
measures to keep the costs at bay and ensure efficiency. One of the most
commonly followed method is popularly known as storage consolidation. While this
approach may work, it has its own associated risks of putting all the eggs into
one basket. A consolidated platform requires a greater degree of data protection
and disaster resilience.

Advertisment

SYNCHRONOUS Vs ASYNCHRONOUS REPLICATION

Synchronous replication used is an in-region hot site that is used for
business continuity and data protection. However, synchronous replication is
limited to relatively short distances- typically less than 50 miles-and is
suitable for replication to in-region recovery sites. This approach does not
protect from regional disasters that may affect both the production site and the
in-region recovery site.

On the other hand, asynchronous replication is used for typically those
organizations that maintain extremely current data copies at out-of-the region
recovery sites. But requirements have evolved over time and proliferating
datasets requiring replication have pushed the limits out of asynchronous
replication solutions. Furthermore, regulations worldwide are endorsing (if not
requiring) out-of-region replication for critical industries such as banking and
securities trading. A number of issues have arisen that indicate the need for an
improved solution.

One of the biggest issues that arise out of using remote copy solutions is
that they consume tremendous amounts of resources. In storage-based solutions,
replication uses part of the storage system cache to capture changes and
transmit them to the other side. It also uses processing cycles on the storage
systems - primarily at the originating (production) data center. These
resources are, in effect, taken away from production applications. The result is
lower application performance, or the increased cost of adding resources to
maintain required performance and throughput. And with data growing
exponentially, this problem is further aggravated. In such a situation, the
obvious solution would be to return the IT resources to where they belong —
the applications.

Advertisment

While both synchronous and asynchronous remote replication processes can
co-exist within an organisation, existing solutions require storage for multiple
copies of the data, as well as complex management and scripting.

DISK-BASED JOURNALING COST-EFFECTIVE

In such a situation, a replication strategy that uses a disk-based
journaling and a pull-based replication engine to reduce resource consumption
and costs, may turn out to be the best bet. A replication solution providing
these features can make data protection and business continuity more efficient
and cost-effective than traditional replication methods.

Using this kind of a strategy would mean that the replication solution would
essentially, while collecting the data to be replicated, write the designated
records to a set of journals volumes. By writing the records to journal disks
instead of keeping them in cache, the replication solution overcomes the
limitations of earlier asynchronous replication methods. Further in the process,
writes to the journal can be cached for application performance reasons, but
they can be quickly de-staged to disk to minimize cache usage. In order to
achieve this, the journal disks have to be specially architected and optimized
for maximum performance.

Advertisment

In order to ensure guaranteed data integrity, the journals can contain
metadata for each record to ensure the integrity and consistency of the
replication process. Each transmitted record set should include a time stamp and
sequence number information, enabling the replication engine to verify that all
the records are received at the remote site, and to arrange them in the correct
write order for storage.

Coupled with the disk-based journaling strategy, use of a pull-based
replication process can create one of the most effective remote replication
solution. This approach would have a remote replication engine to pull the data
from the primary storage system's journal volume across the link. And write it
to the journal volume at the receiving site. The replication engine then applies
the journaled writes to the remote data volumes, using metadata and consistency
algorithms to ensure data integrity. Since the engine that controls asynchronous
replication is located on the remote system, this approach shifts most of the
replication workload to the remote site, reducing resource consumption on the
primary storage system and improving production application performance. In
effect, this kind of an approach restores primary site storage to its intended
role as a transaction processing resource, not a replication engine.

PULL-BASED REPLICATION

By using local disk-based journaling and a pull-based remote replication
engine, the solution releases critical resources that are consumed by other
asynchronous replication approaches at the primary site, such as disk array
cache in storage-based solutions, or server memory in host-based software
approaches. This kind of a solution improves cache utilization, lowering costs
and improving performance of production transaction applications. It also
maximizes the use of bandwidth by better handling the variations of the
replication network resources, enabling enterprises to manage bandwidth cost and
RPO more flexibly and intelligently.

Advertisment

The pull-based replication engine also contributes to resource optimization.
It controls the replication process from the secondary system and frees up
valuable production resources on the primary system. Such a solution can also
increase resilience if the replication solution logs the changes to the journal
disk at the primary site. It should also update the data in the secondary site
without any loss of currency of the data in case of a network or bandwidth
outage. Further, if the replication solution is able to pull data depending upon
the available bandwidth by buffering journal volumes at the primary site when
there is no adequate bandwidth available for transfer, it can result in vastly
improved RTOs and RPOs coupled with lower costs and increased resilience.
Moreover, if the solution enables mapping this kind of an approach across 3 Data
Center configurations, organizations can benefit significantly out of a more
efficient, affordable and cost effective solution for their data protection
needs.

Lim Beng Lay is Product
Manager, Asia-South at Hitachi Data Systems