Leadership

The bus factor has a literature

Knowledge concentration can be measured from the repository itself, and researchers have been doing it for a decade. On the number nobody computes until the resignation.

Code, Noted2 min readLeadership

Every engineering organization tells bus-factor jokes and almost none computes the number. This is odd, because the number is computable, has been for a decade, and the computation runs on data every team already owns: the version history.

Engineering drawing of a structure supported by many columns with load traced to only two

The research literature took the joke seriously before industry did. The truck-factor estimation work of Avelino and colleagues formalized the intuition: mine the repository for file authorship, weight by degree of ownership, and find the smallest set of developers whose departure leaves most of the codebase without a knowledgeable owner. Their survey of popular open-source projects found what practitioners suspected and preferred not to quantify: a large share of significant projects sat at a truck factor of one or two. Subsequent replications have not been kinder.

Two years ago this publication argued that retirements are architecture events only in the sense that all authority questions are; the more direct statement is simpler. Knowledge concentration is a structural property of a system, as real as coupling, and unlike most structural properties it can be measured from artifacts already in hand, in an afternoon, with open tooling. The measurement is imperfect: authorship is a proxy for understanding, the departed sometimes documented well, the present sometimes understand nothing they committed. Proxies with known errors still beat folklore with unknown ones.

What distinguishes the organizations that act on the number is what they treat it as: a maintenance metric, not an HR one. A subsystem at factor one is a finding, like a missing backup; the responses are the ordinary ones, pairing, documented walkthroughs, review rotation, and the metric re-run quarterly shows whether they worked. The failure mode is measuring people instead of systems, which converts a structural diagnostic into a performance conversation and teaches everyone to game authorship immediately.

The essay-sized conclusion: the industry's most cited risk joke has been quietly promoted to a computable quantity, the computation costs nothing, and the only organizations that run it are, by observed correlation, the ones that least need to. Run it before the resignation letter does it for you, at retail, with interest.