Jake Lawrence · Mining Software Repositories / Library Science / Knowledge Graphs · Field Data theme
Every merged pull request in this repository's own history, audited, tagged, and knowledge-graphed, using the same content-as-code method the site turns on everything else, turned on the site itself. On top of it, hand-written case studies about how the site gets built, each cited into that corpus and receipt-checked in the browser.
A field-data audit of this repository's own engineering history: 2,040 of its 2,071 merged pull requests, classified along two independent facets (a controlled-vocabulary surface and a cross-cutting discipline, never a single flat tag) by a deterministic title-keyword ruleset, no runtime model and no per-PR diff. The disciplines roll up into a knowledge graph of about twenty nodes, connected by title co-mentions and same-day merges, each edge traceable to the rule that produced it. The build itself surfaced a finding the prompt did not ask for: thirty-one pull requests describing a now-removed employer-internal admin tool, caught only by reading actual titles against the site's standing red-line policy and withheld, counted but never named. A multidisciplinary literature review across mining software repositories, library and information science, knowledge-graph construction, case-based reasoning, and network science grounds the tagging and graph design in real, cited work rather than invented categories, and the page includes its own build as a case: this is a case study of building a case study library.
Classification as Infrastructure argues every planning layer is a classification system, deliberate or accidental. The Case Study Library builds one on purpose, on the repository's own history, and the surprise is what a supposedly neutral facet caught anyway: an employer's internal tool, invisible until the tags were read closely enough to name it.
Two knowledge graphs built to stay legible rather than impressive. One Load-Bearing Company keeps twenty-six nodes so a single filing's concentration stays readable; the Case Study Library keeps about twenty disciplines so two thousand pull requests do too, and both refuse the bigger graph that would explain less.
Two audits turned on their own author. Italian as Black measures 902 of the author's own chess games for an unexamined habit; the Case Study Library measures 2,071 of his own pull requests the same way, and both find the pattern only after aggregating what no single instance would show.
This is field data, not commentary: dated, sourced, and versioned. Subscribe and the next release, correction or new thread reaches you when it lands. It is the one list for the whole site, so you will get the rest of the work too. No schedule, easy out.
Back to the release table · See it on the network · Machine-readable catalog
I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.
Work with me →