Mainframe Path Start learning free
Applied10 min readLesson 2 of 3

Replication, change data capture and analytics access

Many systems need mainframe data without changing where it lives. Change data capture copies committed changes as they happen; virtualization lets tools query the data in place; streaming and data lakes feed analytics. Each trades freshness, cost and load differently.

Why make mainframe data available

Analysts, data scientists and digital channels want customer and transaction data quickly and in familiar formats. Historically this meant nightly extract jobs that wrote files, converted them and sent them on, so data was a day old. Modern approaches aim to deliver it sooner while keeping the mainframe as the place where the data is actually maintained.

Change data capture

Change data capture (CDC) reads committed changes, usually from the database's recovery log, and sends them to a target such as another database, a message queue or a streaming platform. Because it reads the log, it adds little load to the application and needs no change to the COBOL programs.

Log-based change data capture
ApplicationCOBOL updates Db2 or IMS
Recovery logcommitted changes
Capturereads and filters
Apply or publishtarget DB, MQ, Kafka

Options for analytics access

ApproachHow it worksFreshnessWatch out for
Batch extractScheduled jobs unload data to filesHours to a dayOld data; conversion errors in file transfer
Replication (CDC)Changes copied continuously to a target databaseSeconds to minutesTarget schema drift; monitoring lag
VirtualizationQueries run against the data where it lives through a virtual layerLiveQuery load on production; cost of processor time
StreamingChange events published to a platform such as Apache KafkaSecondsEvent ordering, schema management, consumers falling behind
Data lake or warehouseData lands in a central analytics store, often fed by CDC or streamingDepends on feedGovernance and copies multiplying

Data virtualization (for example IBM Data Virtualization Manager for z/OS) presents VSAM, IMS, Db2 and other sources as if they were tables, so tools can query them with SQL without copying. It gives the freshest data but runs the work on the mainframe, so heavy analytical queries need control, often through WLM classification and offloading eligible work to zIIP processors where the product supports it.

Choosing an approach

Operating a CDC feed

A CDC pipeline is a production service. Teams monitor latency (how far the target is behind), errors on apply, and log retention: if capture stops for longer than the logs are retained (active and archive), it cannot catch up from the log and the target must be reloaded. Schema changes on the source, such as a new column, must be coordinated with the replication configuration.

Replication monitor summary (illustrative)
Subscription  CUSTFEED    State  ACTIVE
  Source      DB2P.CUSTOMER
  Target      kafka topic customer.changes
  Latency     4 sec      Rows applied today  1,284,330
  Errors      0          Last log position read  12:02:41

Common mistakes

Treating a replica as a second place to update

If both copies are updated, they diverge. Replicas fed by CDC should be read-only.

Running unrestricted analytics through virtualization

Large queries consume production processor time. Classify the work in WLM and agree limits with capacity planning.

Ignoring capture lag and log retention

If capture stops for longer than logs are retained, the target needs a full reload. Monitor latency and alert early.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← From VSAM and IMS to relationalConversion pitfalls, quality, governance and the system of record →