Grounding, verification and failure modes
Assistants give better answers when they are fed your site's own material, which is what retrieval-augmented generation does. Even then they make characteristic mistakes, so every AI-assisted change is verified the same way: compile it, test it, compare outputs and have a person review it.
Why grounding matters
A general model knows public COBOL and JCL, but not your copybooks, naming standards, error-handling conventions or that program ACCT0450 must never be run twice. Its answers about your system are guesses dressed as facts. Grounding means giving the model trustworthy, site-specific material to answer from.
Retrieval-augmented generation in concept
Retrieval-augmented generation (RAG) is a common way to do this. Instead of retraining the model, a retrieval step searches an indexed knowledge base for passages relevant to the question and places them in the prompt. The model then answers using those passages, ideally citing them.
Typical sources on a mainframe team are source code and copybooks, runbooks, standards documents, past incident write-ups and product documentation. The quality of the answer can only be as good as what is retrieved: stale runbooks produce stale answers. Access control matters too — a RAG system must not show people documents they would not otherwise be allowed to read.
Characteristic failure modes
| Failure | What it looks like | How you catch it |
|---|---|---|
| Hallucinated APIs or options | A CICS command option, utility parameter or DB2 function that does not exist | Compile or syntax check; look it up in the product documentation |
| Wrong copybook assumptions | Field lengths, types or positions invented or taken from an old version | Compare against the copybook actually included at compile time |
| EBCDIC and packed-decimal errors | Text where packed data belongs, ASCII sort order, wrong record length | Run against real-format test data; inspect in hex |
| Plausible but wrong logic | A rewrite that handles the common case and drops an edge case | Unit tests plus output comparison against the old version |
| Platform confusion | Advice that is correct for Linux or Java but not for z/OS | Review by someone who knows the platform |
Verification techniques
Verification is not optional polish; it is what turns AI output into something you can put your name to. Use layers, because each one catches different faults.
- Compile: catches invented syntax and many wrong field names. A clean compile proves very little else.
- Run tests: unit tests for logic; integration tests for files, DB2 and CICS interactions.
- Compare outputs: run old and new code on the same input and compare results byte for byte. Differences must be explained, not waved through.
- Human review: a reviewer who did not write the prompt reads the change and the evidence.
AI in code review and modernization
In review, assistants can point out missing FILE STATUS checks, unreachable paragraphs or inconsistent naming. In modernization work they can help explain a program before it is restructured, draft a service interface, or suggest how COBOL logic might look in Java. All of these produce candidates for engineers to judge. Converting a program is easy to start and hard to prove equivalent; output comparison over large, representative test data is how teams build that proof.
Records compared : 125,000 Matched : 124,871 Different : 129 First difference : record 1,042 field INT-AMT old: 0000013.47 new: 0000013.46
Common mistakes
Compilation proves syntax, not behaviour. Run tests and compare outputs before trusting a change.
Retrieval returns whatever is indexed, including outdated runbooks. Check the cited source and its date.
The author of a prompt tends to accept its output. A separate reviewer catches what the author is primed to miss.
What you will see at work
- Some sites run internal assistants grounded on their own standards and runbooks, with access controls on what can be retrieved.
- Refactoring and conversion projects rely on large-scale output comparison as evidence of equivalence.
- Code reviewers ask to see test and comparison evidence for AI-assisted changes, not just the diff.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.