Mainframe Path Start learning free
Applied11 min readLesson 2 of 3

MSU, capping and capacity planning

Mainframe software cost is often tied to processor use measured in MSU, commonly the peak rolling four-hour average per LPAR. Capping limits that peak, at the price of slowing work. Capacity planning uses SMF and RMF trends to predict when the machine or the bill will run out of room.

MSU: the unit cost is measured in

An MSU (millions of service units per hour) is a measure of processor capacity used by IBM and other vendors for software pricing. Each machine model has an MSU rating, and RMF and SMF record how many MSUs each LPAR consumes. For many sites, software licences cost more than the hardware, so MSUs get close attention from management.

The rolling four-hour average

Under IBM's sub-capacity pricing models, monthly charges for many products are based on the peak of the rolling four-hour average (R4HA) of MSU use during the month. For each product, the R4HA values of all the LPARs running it are added together hour by hour, and the highest combined value is what is charged. Every few minutes the average of the last four hours is recalculated; the highest value of the month matters most. IBM's Sub-Capacity Reporting Tool (SCRT) reads SMF types 70 and 89 to produce the report the customer submits.

From usage to the bill under sub-capacity pricing
LPAR uses CPUmeasured in MSUs
SMF 70 and 89recorded every interval
R4HArolling 4-hour average
Monthly peakreported via SCRT
Software billfor eligible products

Two consequences follow. A short spike barely moves a four-hour average, but a long heavy period does. And moving heavy work away from the busiest hours of the month can lower the peak without reducing total work. Pricing models change over time and differ by contract; IBM's Tailored Fit Pricing Enterprise Consumption Solution, for example, charges on total MSU consumption rather than the peak, which changes what is worth optimising. Always check which model your site uses.

Capping

MechanismWhat it doesNotes
Defined capacity (soft cap)Limits an LPAR's R4HA to a set MSU valueWhen the R4HA reaches the limit, the LPAR is held down until it falls
Group capacityShared R4HA limit across a group of LPARsLPARs within the group can borrow unused capacity from each other
Hard capping of LPAR weightLPAR never uses more than its share of the machineSet in the hardware partition definitions; does not use the R4HA
WLM resource groupsLimits CPU for selected service classesControls specific workloads rather than the whole LPAR

A defined capacity protects the bill but can slow everything when the cap is reached. WLM still favours important work under a cap, but low-importance work may almost stop. RMF shows capping in the CPU Activity Partition Data and as CPU delay, even when the physical machine has idle processors.

Capacity planning from trends

Capacity planning asks: when will we run out of processor, storage, I/O or cost budget, and what should we do before then? It starts from history kept in SMF:

  1. Collect peak and average usage per LPAR and per major workload, usually weekly and monthly, from RMF data.
  2. Separate growth from one-off events such as a migration or year-end processing.
  3. Fit a trend and add known business changes: a new channel, an acquisition, a product launch.
  4. Compare the forecast with installed capacity, defined capacity and budget, keeping headroom for peaks and failover.
  5. Decide on actions: tuning, rescheduling, zIIP offload, more capacity, or a pricing model change.

Common mistakes

Watching total CPU but not the R4HA peak

Software cost under peak-based pricing follows the highest four-hour average of the month, so the timing of heavy work matters as much as its size.

Setting a cap and forgetting it

A defined capacity set for last year's workload can silently throttle this year's month end. Review caps with every capacity review.

Planning from averages

Averages hide peaks. Use peak intervals, business calendar events and failover needs.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Response time, throughput and bottlenecksInvestigating performance problems →