Mainframe Path Start learning free
Applied11 min readLesson 2 of 3

Recovery, threadsafe and modern interfaces

CICS protects data by backing out a failed task's changes and recovering in-flight work after a crash. It runs programs on more than one kind of TCB, which is why threadsafe matters, and it moves large data with channels and containers and exposes programs as web services and JSON APIs.

Units of work and backout

Every CICS task runs inside a unit of work (UOW). Changes to recoverable resources are tentative until a syncpoint, which happens at the end of the task or when the program issues EXEC CICS SYNCPOINT. If the task abends first, dynamic transaction backout reverses its changes so the data is left as it was at the last syncpoint.

ResourceHow it becomes recoverable
VSAM fileRECOVERY attribute on the file definition (NONE, BACKOUTONLY or ALL), or for RLS files the LOG attribute in the ICF catalog (NONE, UNDO or ALL)
Temporary storage queueDefined as recoverable through a TSMODEL
Intrapartition transient data queueLogical recovery on the TDQUEUE definition
DB2, IMS, MQCoordinated with CICS through two-phase commit

Non-recoverable resources are not backed out. A common surprise is a work file defined with RECOVERY(NONE): the abend backs out the DB2 update but leaves the file record written.

When the region itself fails

CICS writes before-images and UOW status to its system log (DFHLOG, with DFHSHUNT for long-running or shunted work), held in the z/OS system logger. After a failure, an emergency restart reads the log and backs out every UOW that was in flight. If CICS loses contact with a coordinator during two-phase commit, the UOW is indoubt and may be shunted: its locks are kept until the outcome is known. Forward recovery (rebuilding a damaged file from a backup plus logged changes) uses a separate forward recovery log and a product such as CICS VSAM Recovery.

The open transaction environment and threadsafe

Historically all application code ran on one QR TCB (quasi-reentrant), so only one task executed application code at a time in a region and programs could share storage safely. The open transaction environment adds pools of open TCBs (for example L8 and L9) where code can run in parallel.

Why threadsafe saves CPU
Program not threadsafe
Runs on the QR TCBEach DB2 or MQ call switches to an open TCB and backTwo TCB switches per call add CPUAll such tasks share one TCB
Program threadsafe
Can stay on the open TCB after the first DB2 callFewer TCB switchesWork spreads across many TCBsProgram must protect shared storage itself
Program attributes that control TCB useWhat it means
CONCURRENCY(QUASIRENT)
Always returns to the QR TCB; safe default for old code
CONCURRENCY(THREADSAFE)
Can run on whichever TCB it is on; code must be threadsafe
CONCURRENCY(REQUIRED)
Always runs on an open TCB
API(OPENAPI)
Allowed to use non-CICS APIs on an open TCB

Threadsafe is a property of the code, not a switch. A program that updates a shared area such as the CWA without serialising (for example with EXEC CICS ENQ) can corrupt data once marked threadsafe. Teams review code before changing the attribute.

Channels and containers

The COMMAREA is limited to 32 KB. A channel is a named set of containers, each holding any amount of data subject to storage. Programs use PUT CONTAINER, GET CONTAINER and LINK ... CHANNEL(...). CHAR containers can be converted between code pages by CICS; BIT containers are passed unchanged. Channels are the normal interface for web services and APIs.

Web services, JSON and APIs

The details are covered in the APIs course; here the point is that the same COBOL program can serve 3270, MQ and web clients if its interface is clean.

Common mistakes

Assuming every resource backs out

Only recoverable resources are backed out. Check the RECOVERY setting of files and queues the program updates.

Marking a program threadsafe without reviewing it

Shared storage updated without serialisation can be corrupted when tasks run in parallel. Review the code and test under load first.

Cold starting after a failure

A cold start throws away the information needed to back out in-flight work. Use an emergency or auto start unless the systems programmer decides otherwise.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Regions, topology and resource definitionsPerformance, monitoring and troubleshooting →