Generating unit tests and test data
Assistants can draft unit test cases and synthetic test data for COBOL programs, which saves time on the dull parts of testing. But mainframe data has traps — packed decimal, EBCDIC, fixed record layouts — that models often get wrong, so generated tests and data must be checked before you trust a green result.
Why AI is useful here
Many mainframe programs have few or no automated tests. Writing them by hand means reading the logic, listing the cases, building input records and working out expected outputs. An assistant is good at the first two steps: given a paragraph with a nested EVALUATE, it can list the branches and suggest a case for each, including the boundary values people forget.
On z/OS, unit testing is commonly done with IBM's zUnit (part of the IBM Developer for z/OS family) or with vendor and open-source test frameworks; many teams also run tests in GnuCOBOL off-platform for simple logic. What matters for this lesson is not the framework but the inputs and expectations you feed it.
From logic to test cases
Task: List unit test cases for paragraph 3000-SET-FEE.
Logic: EVALUATE TRUE
WHEN ACCT-TYPE = 'P' AND BAL-AMT LESS THAN 1000 MOVE 5.00 TO FEE-AMT
WHEN ACCT-TYPE = 'P' MOVE 0 TO FEE-AMT
WHEN ACCT-TYPE = 'B' MOVE 12.50 TO FEE-AMT
WHEN OTHER PERFORM 9000-BAD-TYPE
END-EVALUATE
Format: table of ACCT-TYPE, BAL-AMT, expected FEE-AMT or outcome.
Include boundaries and invalid values.A good answer includes a balance of 999.99 and exactly 1000.00, a business account, an invalid type such as a space, and a negative balance. You still decide whether those expectations match the business rule — the model only knows the code, and the code may itself be wrong.
Synthetic test data: never production
Test data should be synthetic: made up to have the right shape and edge cases, with no real customer information. Copying production records into a prompt, or asking an assistant to 'make data like this' from a real extract, breaks data-protection rules at most organisations. Where realistic data is needed, sites use approved masking or synthetic-data processes, not ad hoc AI prompts.
The mainframe data traps
Models learn mostly from code written for ASCII systems with variable-length text and floating-point numbers. Mainframe records are different, and these differences produce test data that looks right in a chat window and fails on z/OS.
05 CUST-ID PIC X(8).05 BAL-AMT PIC S9(7)V99 COMP-3.05 OPEN-DATE PIC 9(8).05 FILLER PIC X(5).- Packed decimal (COMP-3): a value like 1234.56 is stored as hex digits with a sign nibble, not as the characters '1234.56'. Text typed into a file is not valid packed data, and invalid packed data typically causes an S0C7 data exception at run time.
- EBCDIC: the same character has a different code from ASCII ('A' is X'C1' in EBCDIC). Data built on a laptop needs a correct code page conversion — and packed or binary fields must not be converted as text.
- Sort order: in EBCDIC, lowercase letters sort before uppercase and digits sort last; in ASCII digits come first. Expected outputs for sorted reports differ.
- Fixed lengths: the record must match the copybook's total length and the dataset's LRECL exactly.
Checking that the tests are real tests
- Compile and run the generated tests; fix anything that does not build.
- Make a test fail on purpose (change an expected value) to prove it can fail.
- Compare outputs against a golden output from the current program before changing anything.
- Review expectations with someone who knows the business rule.
- Commit the tests with the code so they run in the pipeline.
A field is PIC S9(5)V99 COMP-3. How many bytes does it occupy?
Show a hint
7 digits plus a sign nibble, two nibbles per byte.
Show the solution
7 digits + 1 sign nibble = 8 nibbles = 4 bytes.
Common mistakes
A test file with '1234.56' where COMP-3 is expected causes S0C7 or wrong results. Generate records with a tool that understands the copybook.
Generated tests may assert nothing useful. Break an expectation deliberately to prove the test can catch a fault.
Production data in prompts breaches policy and often law. Use synthetic or approved masked data only.
What you will see at work
- Teams adding tests to legacy programs use assistants to list cases, then build data with existing utilities.
- Before a refactor, the current program's output is captured as a baseline and new versions are compared against it.
- Test data requests go through data-management processes, not chat windows.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.