Call graphs, data lineage and impact analysis
Once you have an inventory, you relate the pieces: which programs call which, which jobs and programs read and write which data, and what a proposed change will touch. Those relationships drive impact analysis for every change.
Call graphs
A call graph shows which programs invoke which others. In COBOL batch, the main mechanism is the CALL statement. In CICS, programs also transfer control with LINK and XCTL, and start other transactions. A graph built from these shows the full chain behind a job step or a transaction ID.
* Static: the name is a literal, visible to analysis
CALL 'ACCTVAL' USING WS-ACCOUNT-REC.
* Dynamic: the name is in a variable, set at run time
MOVE 'RATE' TO WS-PGM-PREFIX.
MOVE WS-PRODUCT-CODE TO WS-PGM-SUFFIX.
CALL WS-PGM-NAME USING WS-RATE-REC.The first call is easy for any tool to find. The second is the classic blind spot: the program name is assembled from data, so static analysis sees only that some program is called. Good tools try to resolve the possible values; where they cannot, you need runtime evidence or a developer who knows the product codes.
Data lineage
Data lineage traces where data comes from, what changes it and where it goes. On the mainframe this means following records through jobs, datasets, programs, Db2 tables and interfaces. Typical questions: which feed populates this field in the statement file, and which downstream systems receive it?
In batch, the JCL tells you which datasets each step reads and writes, through the DD statements and their DISP. Inside each program, the file definitions and the copybook for the record layout tell you which fields are involved. GDGs, temporary datasets and datasets whose names are built from symbols all need care, because the name in the JCL is not always the name at run time.
Impact analysis
Impact analysis answers: if I change this, what else is affected? It is the most common everyday use of discovery data, long after the modernization project has finished.
| Proposed change | What to trace | Typical findings |
|---|---|---|
| Widen a field in a copybook | Every program that copies it, every file and table that stores the record | Recompile list, file conversions, downstream extracts with fixed layouts |
| Add a column to a Db2 table | Packages that use the table, SELECT * usage, unload jobs | Programs that need rebind or change, extracts that need new layouts |
| Retire a batch job | Datasets it creates and who reads them | Downstream jobs or external transfers that would starve |
| Change a called subprogram | All callers, static and dynamic | Batch and online callers with different expectations |
Working with tool output
- Treat graphs as hypotheses. Confirm surprising links, and missing links where you expect them, with source or runtime data.
- Record unresolved dynamic calls and dataset names explicitly, so nobody mistakes 'not found' for 'not used'.
- Note links that leave the application: MQ queues, file transfers, APIs and shared tables owned by other teams.
- Keep the repository refreshed from the source control system, or impact analysis will drift out of date.
Common mistakes
A CALL using a variable name is invisible to simple text search. Resolve the possible names from data or runtime evidence, and list what remains unresolved.
A layout change also affects files, tables and every downstream consumer, including external partners. Follow the data, not only the code.
Names appear in comments, in symbols and in built-up strings. Use parser-based tools where available and verify the results.
What you will see at work
- Change requests for regulated systems often require an impact analysis section listing affected programs, jobs and interfaces.
- Data governance and audit teams ask for lineage of key figures such as balances or risk measures.
- Developers use call graphs to understand an unfamiliar transaction before fixing a defect.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.