The four VSAM types
VSAM is a family of file organisations for data you need to reach quickly by key or position, rather than by reading from the start. Most of what you meet in practice is one of the four types below, and usually the first one.
Vocabulary you need before anything else
| Term | Meaning |
|---|---|
| Cluster | The whole VSAM object — what you name and refer to |
| Data component | Where the records live |
| Index component | For a KSDS: the structure that maps a key to a location |
| Control interval (CI) | The unit VSAM reads and writes — like a block |
| Control area (CA) | A group of control intervals |
| Free space | Room deliberately left empty inside CIs and CAs so inserts do not force a split |
| Alternate index (AIX) | A second way to reach records, by a non-primary key |
| Path | The object a program opens to read through an alternate index |
How a KSDS finds a record
The index is a tree. The top levels — the index set — narrow the search to a control area; the bottom level — the sequence set — points at the exact control interval. This is why lookup cost barely grows as the file gets bigger, and why the key must be unique and fixed in position and length.
Choosing between VSAM and DB2
| Use VSAM when | Use DB2 when |
|---|---|
| Access is always by one known key | You need flexible queries and joins |
| The record layout is stable and owned by one application | Many applications and tools need access |
| You need the lowest possible access cost | You need SQL, reporting, and referential integrity |
| The application already uses it | You are designing something new |
Common mistakes
It is a file organisation. There is no SQL, no query optimiser and no cross-file referential integrity.
The primary key's position and length are fixed at define time. Changing it means defining a new cluster and reloading.
That is what browse commands are for, and they should be bounded. Unbounded browses online are a classic performance defect.
What you will see at work
- Many core applications are VSAM-based and will remain so. It is not a legacy detail; it is where the data is.
- Alternate indexes must be kept in step with the base cluster. When they drift, lookups return wrong or missing records.
- The names of the data and index components usually follow a site pattern such as .DATA and .INDEX suffixes.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.