Automated testing on z/OS
A pipeline is only as safe as its tests. On the mainframe that means unit tests for individual COBOL programs, integration tests against test copies of DB2, CICS and files, and regression checks that compare today's output with a known-good run.
Three layers of tests
- Unit tests run a single program with controlled inputs and check its outputs. Frameworks such as IBM's zUnit can stub out file, DB2 and CICS calls, so the test needs no shared environment.
- Integration tests run the program for real against test databases and regions. They catch problems unit tests cannot, such as a missing DB2 bind or a wrong dataset name.
- Regression tests run a batch job and compare its output files with a saved, approved copy. Any unexpected difference fails the stage.
Test data is the hard part
Production data cannot simply be copied into test: it contains personal and financial details. Teams build small, synthetic data sets, or use masking tools that replace real names and account numbers while keeping the formats valid.
Where tests run
Some teams run tests on a dedicated test LPAR; others use emulated z/OS environments for developers and early stages, keeping the shared test LPAR for integration. Either way, each pipeline run should start from a known state, so results are repeatable.
| Stage | Typical check | Fails when |
|---|---|---|
| Build | Compile and link | RC 8 or higher |
| Unit test | zUnit or similar | Any assertion fails |
| Integration | Run against test DB2/CICS | Abend, wrong SQLCODE, wrong output |
| Regression | Compare batch output | Unexpected differences |
Common mistakes
Most production failures come from bad input: spaces in numeric fields, empty files, duplicate keys. Write tests for those first.
If another team rewrites the test database, your tests fail randomly. Load a known data set at the start of each run.
This breaks privacy rules in most countries. Use synthetic or masked data.
What you will see at work
- Coverage of legacy code is usually low. Adding tests around the code you change is a realistic goal; testing everything at once is not.
- A regression comparison that flags a difference is not always a bug — it may be an intended change. Updating the approved baseline is part of the work.
- Test results attached to a pull request are often what convinces change boards to approve faster releases.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.