When I moved everything over to Forgejo, I mentioned there was still a massive reorganisation of everything to do.

One of the things that has existed for years is my dockerfiles repo. It started out fairly sensibly (well, that's what I tell myself anyway) and eventually became the place where basically every container I run lived. Each directory had a Dockerfile; some had patches, some had bits of shell, some had templates generating other bits of shell, and over time I had accumulated a fairly impressive collection of different ways of saying "download this thing, compile it, and put the resulting binary in an Alpine container".

This worked perfectly well, and I could have carried on using it indefinitely. Dockerfiles and templating just didn't feel ideal, and I definitely felt like I was abusing them. Unfortunately, "it works perfectly well" has never been a particularly effective defence against me deciding to replace something.

I'd looked at melange and apko before, thought they looked interesting, and thought they fixed some of my annoyances and made things a lot more readable, but they had some annoying flaws of their own. I then mostly ignored them because replacing a perfectly functional Docker build setup with two new tools and a package repository would have been difficult to justify even by my standards.

The recent move to Forgejo made me look at them again because I now control quite a lot more of the surrounding infrastructure. Forgejo is doing the git hosting, I can run the CI how I want, organisations are cheap and useful rather than something I need to work around, and the package registry isn't just a place to throw container images. It also functions as a repository for other types of software, including Alpine.

Given my earlier decision to give up on melange and apko was heavily influenced by a lack of alpine registry, obviously it was time to reconsider things... throw everything away and start again. Breaking down what these images actually did: Build some software, bundle it together in a container, host it ready to run.

Software

melange builds APKs, the package format used by Alpine. A melange config describes a package: where its source comes from, what is needed to build it, what it depends on at runtime, and the steps needed to turn that source into the files which belong in the package.

melange is pipeline-based, which I quite like. There are built-in pipelines for the boring common things: fetching source, checking out git, applying patches, building Go programs, and the usual autotools nonsense. Package definitions can compose those rather than every package containing its own slightly different shell script. As I like consistency in my recipes, you can also add your own reusable pipelines. I now have some custom melange pipelines for the bits I kept finding myself repeating across packages, which means the package configs themselves can mostly describe the differences between applications rather than reproducing the same lump of build machinery each time.

Tailscale, for example, needs a runtime directory and the usual licence handling. The package definition doesn't contain shell for either of those things; it just says:

- uses: dirs
  with:
    dirs: ${{vars.runtime-dirs}}

- uses: go/licenses
  with:
    checks: true


dirs knows how I want package directories created, go/licenses knows what I want done with Go licences, and another pipeline, test/paths, handles the equally exciting job of checking that things ended up where I expected them. Quite a few packages need some combination of these, so there wasn't much value in each of them having its own implementation.

This is probably just templating with a more respectable name, but at least this time the templating system is part of the packaging tool rather than something someone wrote to generate Dockerfiles.

The patch handling is another part I particularly like being a little more explicit. A package can include patches alongside its build definition, and applying them is an explicit part of the build pipeline. Fetch the upstream source, verify it, apply the patch, build the result.

Previously, a patch belonged, conceptually at least, to a container. This never felt quite right: the patch isn't a modification to the container. It's a modification to the software.

If three containers use the same piece of patched software, I shouldn't have three containers each knowing how to produce that software. I should have one package which knows how to produce it and three containers which install it.

I got around this previously by having images build bits of software and then using them in multi-stage builds to get what I wanted. It worked, but it meant the container build still owned the process of producing the software rather than simply describing what should be in the container.

My previous experiments with melange did show that there were some issues I had with it which, whilst not fatal, made certain things a little awkward. But a nice package definition that supports patches means I can just use this to patch melange too, and then build everything else with my patched melange.

And that is, in fact, what happens. The package currently applies a collection of patches. Most of these patches just change how it interacts with Forgejo's implementation of an Alpine repository, but they also enable a set of custom update types so my updater can pull updates from sources other than those melange supports. One patches adds the custom pipelines mentioned above, because apparently my build process wasn't convoluted enough, it seemed better to bundle them into the package rather than refer to them separately.

I don't really want to maintain forks of build tooling forever, because maintaining your own patches to the tool you use to maintain your own patches is the sort of recursion which eventually leads to your fork drifting out of date with upstream, but the upstream doesn't seem particularly keen on external patches. So a set of simple patches felt like the next best plan; they may occasionally fail to apply, but hopefully they should be easy enough to fix or drop. I suspect this distinction will become increasingly important to me right up until one of my patches stops applying after an update and I swear at Past Me for writing this paragraph, but hopefully LLMs will stop that from being the case.

The packages themselves get built in Forgejo CI and pushed into Forgejo's Alpine package registry, so I now have my own APK repository containing the relatively small collection of software I actually care about.

For something which already exists in Alpine, I use Alpine's package. For something which doesn't exist there, or where I need a different version or some local changes, I can build my own. I am not attempting to rebuild Alpine. This is mostly a message to myself; I am already being tempted to start packaging busybox and a certificates bundle... I have considered contributing to Alpine to keep packages more up to date, but mostly they lag behind for good reasons.

Containers

apko builds OCI images. The apko configuration says which repositories to use, which packages should be installed, which user the resulting container should run as, what the entrypoint is, environment variables, a few bits of filesystem setup, and other image metadata. If you want some software in the image, it needs to be a package. This initially sounds limiting, and it is. That's the point.

This is, perhaps surprisingly, a much better abstraction for running software. Dockerfiles have always been slightly weird for this. They're excellent at describing the process by which you arrived at a filesystem. That is also their biggest problem.

apko also didn't escape some patches, most of these were because I didn't like the references to apko and Chainguard, but it also needed a minor change to cope with Forgejo as part of the melange patching.

I don't care about all the cruft around compiling and installing the software; I just care that it was done. The fact that at some point halfway through building it I ran rm -rf /var/cache/apk/* is not meaningful information about the software. It's housekeeping caused by the build mechanism leaking into the description of the result.

Multi-stage Dockerfiles make this considerably less terrible, but they're still fundamentally a very elaborate way of building a filesystem by running shell commands and then throwing away the bits you didn't want.

Tailscale is a reasonable example because building the package isn't completely trivial. The package builds three Go binaries — tailscaled, tailscale, and containerboot — creates its runtime directory, and sorts out the licence information.

The container doesn't know any of that. Once the package exists, the interesting end of its configuration is basically:

paths:
  - path: /var/lib/tailscale
    type: directory

environment:
  TS_SOCKET: /var/run/tailscale/tailscaled.sock
  TS_STATE_DIR: /var/lib/tailscale
  TS_TAILSCALED_EXTRA_ARGS: --tun=tailscale0,userspace-networking

entrypoint:
  command: /usr/bin/containerboot


There is rather more machinery involved in producing the Tailscale package than there is in producing that container. That's exactly the direction I want the complexity to flow.

The actual container definitions have consequently become extremely boring. This is good. A container should be boring. The interesting part is what you're using the container for.

Package software as packages. Install packages into a filesystem. Run the filesystem. For example, Miniflux went from a single Dockerfile to separate package and container definitions.

Forgejo

The old dockerfiles repository had everything together. The new setup has an organisation containing repositories for individual packages, and another set of repositories describing the containers that use them. This feels like a small distinction, but it feels a lot more natural in my head.

The source for a package is in a repository in Forgejo. Forgejo runs the CI which builds it. The resulting package is published in Forgejo. The source for the image is in a repository in Forgejo. Forgejo runs the CI which pulls packages from Forgejo as needed, builds the image, and pushes it back into Forgejo's container registry.

There was one thing I didn't particularly want to lose, though, which was the monorepo view. Having every package and container in its own repository is much nicer for actually maintaining them, but it is slightly annoying if you just want to browse everything, grep across the lot, or point someone at one URL and say "all of the stupid stuff is in here".

So the old dockerfiles repository still exists, except it isn't really the source repository anymore. It's a read-only amalgamation of the real repositories in Forgejo. The package repositories get collected under packages/, the container repositories under containers/, and the result is periodically rebuilt as a generated monorepo view, again using Forgejo.

As I mirror a lot of these repos back to GitHub, I now have a backup of the package and container definitions that I could pull and rebuild from without having to dig through backups.

Result

I now have more repositories, more CI jobs, an Alpine package repository I didn't have before, custom melange pipelines, patches to both of the tools doing the building, and two build tools where previously I had Docker. On paper, this sounds like I have made everything substantially more complicated. In practice, it feels simpler.

In reality, I suspect I've technically made myself a very small Linux distribution. Just one with almost no packages, no installer, no kernel, a slightly patched toolchain, some packaging conventions of its own, two different repository layouts depending on how you look at it, and no intention whatsoever of becoming a Linux distribution.

Which is probably how these things start.