Copy services, storage problems and the storage role
Modern disk subsystems can copy and mirror data in hardware, which underpins fast backups and disaster recovery. Most day-to-day storage incidents are about space, recall delays and policy mistakes, and the storage administrator's job is to prevent and fix them.
Copy services in plain terms
Enterprise disk subsystems such as IBM DS8000 provide copy functions inside the storage hardware itself. Other vendors offer compatible or equivalent functions. You do not need to configure them as a beginner, but you will hear the names constantly.
| Function | What it does | Typical use |
|---|---|---|
| FlashCopy | A point-in-time copy that is usable almost immediately; the hardware copies the data in the background | Fast backups, cloning test data, consistent copies of DB2 or IMS |
| Metro Mirror | Synchronous mirroring: each write is completed at the second site before the application is told it finished | Disaster recovery within metro distances, with no data loss |
| Global Mirror | Asynchronous mirroring over long distances, with consistent copies formed periodically | Disaster recovery to a distant site, accepting a small data loss window |
| z/OS Global Mirror | Older asynchronous mirroring driven by z/OS (also called XRC) | Still found at some long-established sites |
FlashCopy is often invoked through DFSMSdss (for example a COPY or DUMP that uses it underneath) or through products such as DB2 and DFSMShsm fast replication. Mirroring is usually managed by automation such as IBM GDPS, which handles switching to the other site in a controlled way.
Common storage problems
| Symptom | Usual cause | First checks |
|---|---|---|
| SB37, SD37 or SE37 abend | Dataset ran out of space or extents | SPACE and secondary allocation in JCL, data class defaults; see x37 abends |
| Allocation fails for a new dataset | Storage group full or volumes disabled for new allocation | Storage group space and volume status with the storage team |
| Job waits a long time at start | Recall from ML2 or a tape mount pending | Job log and system log for recall or mount messages |
| Dataset has unexpected attributes or location | ACS routine assigned different constructs | Constructs shown in ISPF 3.4 or LISTCAT |
| Old data disappeared | Management class expired it | Management class rules, HSM backup versions |
Operators and storage staff can check SMS status from the console. The D SMS,STORGRP(name),LISTVOL command shows the storage group and the status of each of its volumes, and D SMS,VOL(volser) shows one volume. Space and threshold reporting is usually done with ISMF, vendor storage reporting tools or SMF-based reports.
STORGRP SGPROD TYPE POOL VOLUME STATUS SPACE USED PRD101 ENABLED 88% PRD102 ENABLED 91% PRD103 DISNEW 97% PRD104 ENABLED 86%
Production scenario
The storage administrator role
- Designs and maintains SMS constructs, storage groups and ACS routines.
- Runs and tunes DFSMShsm, its cycles and its control datasets.
- Manages tape libraries, virtual tape and the tape management system.
- Plans capacity, adds and removes volumes, and moves data with DFSMSdss.
- Works with disaster-recovery teams on copy services and recovery tests.
- Answers incidents about space, recall delays and lost data.
It is a specialist systems programming role. Many people move into it from operations or production support, where they first learned to read space abends and recall messages.
Common mistakes
Mirrors faithfully copy deletes and corruption. You still need point-in-time copies and backups for logical recovery.
Oversized allocations waste pool space and can fail when no single volume has that much free. Size primary realistically, give sensible secondary space, and consider data class options with the storage team.
If several jobs fail allocation at once, check storage group space and volume status before changing JCL.
What you will see at work
- Space abends are among the most common overnight failures, so knowing SPACE, extents and storage groups pays off immediately.
- Disaster-recovery tests rely on FlashCopy and mirroring; application teams are asked to validate data at the recovery site.
- Storage administrators are usually a separate team, and a clear ticket (dataset name, job, time, message) gets faster help.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.