Birdkeeping

We need new ways of talking about large language models. The artificial intelligentsia at frontier labs have beguiled us (and themselves) into using an exclusively calamitous tone of voice. If we are to take the Silicon Valley oracles at their word, the latest models are effectively sentient and thus pose existential risk, that open weights models are as or more dangerous than nuclear or biological weapons, and that the best thing we can do for humanity’s and the world’s increasingly unstable climate right now is to concrete over open green space with massive data centre developments.

Even among the more technically astute, something seems to have gone linguistically awry. Machine learning researchers compare their LLM hackery to the life sciences and quantum physics, as if the terabytes (petabytes? exabytes?) of numbers operationalised in unimaginably large data centres bear more than a passing resemblance to the atom bomb or the human body. An agent orchestration system called gastown, a kind of tool that many if not most software developers now use when coding, developed by an industry veteran, likens agential development to the Wild West and rhetorically reduces agents to low cost workers in a factory. (Gastown has more recently been upgraded into gascity, a framework that markets itself through the nauseatingly capitalist byline: “Your software deserves a factory”.) A respected VC investor and market analyst suggests thinking about LLM chatbots as a shortcut to infinite interns.

Let’s take a deep breath, and think sensibly about this language for a moment. If we metaphorically manage LLMs schizophrenically as both our workers and managers, our interns and our bosses, our slaves and our masters, all at once, their miraculous but jagged ability to comprehend and reproduce language is abased and apotheosised, all at once. We end up talking about robot rights and relitigating Pascal’s wager, rather than exploring the real potentiality and potentially ethical possibilities of working with LLMs.

Re-routing the flight path of AI discussion

I can understand why thinking about LLMs simply as machines leaves a little bit wanting. LLMs’ uncanny ability to comprehend written text and execute real virtual tasks on a computer when strapped into a coding harness isn’t something we’re used to thinking about as mechanical. Their actions are often not reproducible, and they can abide by the letter of your law while utterly ignoring its spirit. I recently referred to my ‘agent orchestration workflow’—the going industry jargon to denote a specific configuration of LLM use to assist your digital work—as a toolkit for taming jagged intelligence, because I think there is something animalistic about LLMs. Perhaps this something is that, like animals, we are vain to consider that they are actually ever under our power. As Gerard Wajcman argues, we project a kind of perfect world onto animals of many kinds as an imaginative escape from the travesty of humanity.1 It is a fantasy to think that LLMs will obey our commands perfectly once models get better, as language is an imprecise instrument. If LLMs exhibit anything that deserves to be called ‘intelligence’, it is of an animalistic variety: effectively opaque, imperfectly captured by anthropomorphised vocabulary.

With this in mind, I want to throw another metaphor experimentally into the oversized soup of AI takes. What if we are better off thinking about LLMs as birds? Birds can do many things that we humans cannot—deliver messages by air, glide for hours by locking their wings, fly, etc. There are also many things that humans can do that birds cannot. We are not, however, imminently at risk of a bird uprising that will result in humanity’s extinction (to the best of my knowledge).

There are many different species of birds. Albatrosses can fly for longer periods than other species, and hummingbirds can hover in place through rapid wing movements. Almost all species of bird in my native country, Aotearoa New Zealand, are flightless, as given the lack of mammalian predators (before the arrival of humans on the island), rooting around with long beaks in the New Zealand bush was a surer-footed way to find food than searching the skies.

Rather than thinking about LLMs through the mythology of ‘one model to rule them all’ (the ideology that frontier labs are pushing), I suggest that we’re better to think about how to cultivate a thriving and diverse model ecosystem in which many species of models coexist. Massive models may be useful and even environmentally justifiable for certain tasks, such as finding security holes in critical infrastructure. But smaller models are already sufficient for many of the use cases for LLMs such as codebase reconnaissance and harness engineering. It doesn’t make sense to use a Harpy eagle to deliver a small note when a carrier pigeon would do.

As this metaphor already begins to suggest, analogising LLMs to birds also makes the art of taming them appear as an activity that could be artful and personal, rather than necessarily industrial. What if working effectively with LLMs in our day-to-day lives is more like falconry than producing pins in a factory? We can send an LLM off to do certain kinds of work that we trust it to do well, while being cognisant that its operation is always subject to its own obscure and animalistic impulses, at least to some degree, no matter how ‘well trained’ it is.

From beads to birds

LLMs are excellent at implementing directed workstreams when strapped into coding harnesses. The issue is working out how to give them a structured set of directives so that they don’t exploit any unspoken ambiguity that might lurk beneath the surface of these directives. It is also important to track the work they complete so that we can both evaluate their work at reasonable intervals, and roll back or redirect them when we notice something amiss.

Sanguine use of a version control system (VCS) such as git or jj smooths over many of these concerns. So long as LLMs can package their work into comprehensible revision sets (‘revsets’), a human can reasonably both follow along and steer its course with a good amount of granularity. In order to queue up new tasks for LLMs to have a go at, some kind of issue tracking system also goes a long way. The hitch is that neither git nor jj provides an integrated issue tracking system. A VCS lets you register work that has already been done, but it doesn’t have an opinion on how you line up work that needs to be done.

Beneath the capitalist tropology of gastown and gascity (the AI orchestration systems built by Steve Yegge that industrially adopt and scale agents to try to build software), there rests the germ of an important idea in beads, the queueing mechanism for those bigger, more chaotic systems. As Yegge correctly notes, sprawling Markdown plans are a terribly unstructured way to keep an LLM on track: it can execute more complex activity in a reasonable fashion when given a dependency tree of atomically specified units of work.

Having worked with beads almost daily since November 2025, it is hard to imagine developing software with an LLM without something like it. Beyond allowing me to queue up work for agents, it also acts as a local issue tracking system across every codebase that I’m working on. The set of open beads across all of my projects represents all of the work that I have set out to do. Beads has almost fully replaced GitHub issues as my per-project issue tracking solution.

But the divergent philosophy of software development that beads was built to support has started to grate on my everyday use of it. Beads are not intended to be read by humans, and the beads CLI is not intended to be used by humans. I have been using the excellently named abacus, but even through this better lens, the experience of inspecting beads with human eyes constantly reminds me that they were designed to have a jagged intelligibility. Beads are supposed to be parcellated directives for a factory line of interned agents.

Birds are a conceptual alernative to beads. Let us think of birds as agent-ready units of work, analogous to beads, but designed with human legibility as a first-class citizen. Birds could therefore be a reconceptualisation of how to solve many of the same problems of keeping state while working effectively with LLMs in the software development lifecycle.

The art of falconry with birds

Let’s recap. Imagine that LLMs are not interns, workers, or superintelligences, but birds. These computational creatures can convincingly, stochastically parrot human language in the form of text and (to some extent) audio. We can also trust them to execute certain kinds of digital actions, just as we can trust pigeons that have been trained accordingly to deliver messages tied to their leg, or falcons to assist with raptor rehabilitation. I do not offer this metaphor to raise concerns about whether we should let the LLMs fly free, nor do I think of domesticating biological birds as a practice free from serious ethical concerns and questions. What I do believe is that an analogy between LLMs and birds captures the spiky capabilities of the former in 2026 (their ‘jagged intelligence’) better than an analogy between LLMs and humans.

As I have already suggested, if LLMs are birds, coordinating them to assist with the labour of software development is something like falconry. Instead of a cyber Charles Babbage, a silicon superintendent of the industrial information factory, as Yegge’s gascity positions her, I propose that we reframe the effective software developer in the age of LLMs as a master falconer, an algebraic austringer.

Bibliography

  1. 1“We observe them with admiration, with envy, and also in grief, grief at having treated them so badly and resembling them so little. From nature, language keeps us at bay. Man is a denatured animal. We are animals sick with language.” [1, p.131].