Mainframe Path Start learning free
Applied11 min readLesson 3 of 3

Grounding, verification and failure modes

Assistants give better answers when they are fed your site's own material, which is what retrieval-augmented generation does. Even then they make characteristic mistakes, so every AI-assisted change is verified the same way: compile it, test it, compare outputs and have a person review it.

Why grounding matters

A general model knows public COBOL and JCL, but not your copybooks, naming standards, error-handling conventions or that program ACCT0450 must never be run twice. Its answers about your system are guesses dressed as facts. Grounding means giving the model trustworthy, site-specific material to answer from.

Retrieval-augmented generation in concept

Retrieval-augmented generation (RAG) is a common way to do this. Instead of retraining the model, a retrieval step searches an indexed knowledge base for passages relevant to the question and places them in the prompt. The model then answers using those passages, ideally citing them.

RAG over mainframe knowledge
Questionwhat does ACCT0450 do?
Retrievesearch indexed docs and code
Augmentadd top passages to prompt
Generateanswer with sources
Verifya person checks the sources

Typical sources on a mainframe team are source code and copybooks, runbooks, standards documents, past incident write-ups and product documentation. The quality of the answer can only be as good as what is retrieved: stale runbooks produce stale answers. Access control matters too — a RAG system must not show people documents they would not otherwise be allowed to read.

Characteristic failure modes

FailureWhat it looks likeHow you catch it
Hallucinated APIs or optionsA CICS command option, utility parameter or DB2 function that does not existCompile or syntax check; look it up in the product documentation
Wrong copybook assumptionsField lengths, types or positions invented or taken from an old versionCompare against the copybook actually included at compile time
EBCDIC and packed-decimal errorsText where packed data belongs, ASCII sort order, wrong record lengthRun against real-format test data; inspect in hex
Plausible but wrong logicA rewrite that handles the common case and drops an edge caseUnit tests plus output comparison against the old version
Platform confusionAdvice that is correct for Linux or Java but not for z/OSReview by someone who knows the platform

Verification techniques

Verification is not optional polish; it is what turns AI output into something you can put your name to. Use layers, because each one catches different faults.

Layers of verification
Human reviewindependent reviewer reads change and evidence
Output comparisonold vs new on the same data
Testsunit and integration
Compilesyntax and references

AI in code review and modernization

In review, assistants can point out missing FILE STATUS checks, unreachable paragraphs or inconsistent naming. In modernization work they can help explain a program before it is restructured, draft a service interface, or suggest how COBOL logic might look in Java. All of these produce candidates for engineers to judge. Converting a program is easy to start and hard to prove equivalent; output comparison over large, representative test data is how teams build that proof.

Comparing old and new output (illustrative)
Records compared  : 125,000
Matched           : 124,871
Different         : 129
First difference  : record 1,042  field INT-AMT
  old: 0000013.47   new: 0000013.46

Common mistakes

Stopping at a clean compile

Compilation proves syntax, not behaviour. Run tests and compare outputs before trusting a change.

Assuming RAG answers are always current

Retrieval returns whatever is indexed, including outdated runbooks. Check the cited source and its date.

Letting the same person prompt and approve

The author of a prompt tends to accept its output. A separate reviewer catches what the author is primed to miss.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Generating unit tests and test dataBack to AI-assisted code analysis and testing