Packaging binaries: artifacts, versions and repositories
Packaging turns build outputs into one versioned file, called an artifact, that can be stored, checked and deployed. The rule is build once, deploy many: the same package moves through every environment. Artifacts are kept in an artifact repository, with a version number, a checksum and a record of the commit they came from.
Build once, deploy many
If test and production are each built separately, they can quietly differ: another compiler level, another copybook, another dependency. So a pipeline builds once, packages the result into a single artifact, and deploys *that same file* to test, then to acceptance, then to production. What was tested is, byte for byte, what goes live.
What goes into a package
| Platform | Typical build outputs packaged |
|---|---|
| Java | A JAR or WAR file |
| Linux service | A .tar.gz archive, an RPM or DEB package, or a container image |
| z/OS | Load modules, DBRMs for Db2, CICS definitions, JCL and other members produced by the build |
Alongside the files, a good package carries a manifest: the version, the build number, the Git commit it was built from, and the list of parts with their checksums. Anyone can then answer 'what exactly is in release 1.4.2, and where did it come from?'
$ tar -czf dist/payroll-1.4.2.tar.gz -C build . $ sha256sum dist/payroll-1.4.2.tar.gz > dist/payroll-1.4.2.tar.gz.sha256 $ cat dist/payroll-1.4.2.tar.gz.sha256 9b1f0c…e47a dist/payroll-1.4.2.tar.gz
tar -czf creates (c) a gzip-compressed (z) archive file (f). -C build . means 'take everything from the build folder'. sha256sum calculates a checksum, a fingerprint of the file. If even one bit changes later, the checksum no longer matches, so the deployment step can prove it received exactly what was built.
Version numbers
Many teams use semantic versioning, MAJOR.MINOR.PATCH. Increase PATCH for a fix that changes nothing else (1.4.1 → 1.4.2), MINOR for a backwards-compatible new feature (1.4.2 → 1.5.0), and MAJOR for a change that breaks existing users (1.5.0 → 2.0.0). Pipelines often add the build number or commit ID as well, so that every artifact is unique.
Artifact repositories
Artifacts are stored in an artifact repository, not in Git. Git is built for source code; a binary repository is built for large files that never change once published. Well-known products include JFrog Artifactory and Sonatype Nexus Repository, and GitHub Packages offers something similar. A release version, once published, should never be overwritten. If it is wrong, publish a new version.
Under semantic versioning, which number do you increase for a bug fix that changes nothing else: MAJOR, MINOR or PATCH?
Show a hint
The last of the three.
Show the solution
PATCH: 1.4.1 becomes 1.4.2.
Examples are for learning. Run commands and jobs only on a system you are authorised to use, such as a training or test system, and never on production without approval.
Common mistakes
Separate builds can differ in subtle ways. Build once and promote the same artifact.
If 1.4.2 can change after release, nobody knows which 1.4.2 is running. Publish 1.4.3 instead.
Large binaries bloat every clone and are not what Git is designed for. Put them in an artifact repository.
What you will see at work
- When production misbehaves, the first questions are 'which version is running?' and 'what commit was it built from?' The manifest answers both.
- Release notes are often generated from the commits between two versions.
- Retention rules in the artifact repository decide how far back you can roll back. Check them before you need them.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.