The Librarian's Paste Pot vs. The Coder's Spider: Two Philosophies of Preservation
In the quiet urgency of preserving our digital commons, two distinct camps have emerged, each with their own tools, anxieties, and visions of posterity. They are not so much rivals as they are distant cousins, speaking different dialects of the same preservationist language. On one side stands the tradition of the Librarian’s Paste Pot; on the other, the method of the Coder’s Spider. One is manual, curatorial, and deeply human. The other is automated, comprehensive, and profoundly algorithmic. The friction between them reveals a fundamental tension in how we decide what is worth saving.
The Librarian’s Paste Pot approach is as old as libraries themselves, simply translated to a new medium. Imagine a dedicated individual, perhaps a community historian or a specialized archivist, who identifies a vital but fragile digital resource—a local news blog facing shutdown, a unique digital art project, a disappearing government FAQ. Their method is one of selective salvation. They use tools that are, in essence, sophisticated digital paste pots: wget for recursive downloads, browser extensions for single-page saves, manual organization into folders with clear, meaningful names. This is preservation by hand, guided by expertise and a specific narrative. The result is a curated collection, a hand-picked bouquet of data whose value is determined by a human eye long before the final file is saved.
Conversely, the Coder’s Spider operates on a scale that defies human curation. Its most famous practitioner is the Internet Archive’s Wayback Machine, a vast automaton that sends out digital spiders to crawl the web, attempting to capture everything in its path. There is no initial judgment of value; a corporate homepage is scraped with the same algorithmic indifference as a teenager’s poetry blog. This is preservation by volume, driven by the philosophy that significance is often only apparent in retrospect. The spider doesn’t know what will be important to future historians; it only knows to capture. Its strength is its indiscriminate nature, creating a massive, un-curated fossil record of the digital age.
The core difference lies in the moment of appraisal. The Librarian appraises before preservation, saving only what is deemed significant. The Coder appraises after, saving first and asking questions later. Each method carries a complementary weakness. The Paste Pot is vulnerable to the curator’s blind spots and biases; what seems unimportant today might be a crucial cultural artifact tomorrow. It’s a method that can inadvertently erase the very context it seeks to preserve by being too selective. The Spider, however, captures context in its chaotic entirety, but in doing so creates a haystack of unimaginable proportions, within which the needles of truly important data may be lost forever, or rendered inaccessible by the sheer weight of the captured whole.
Ultimately, we need both. The Spider’s automated, sprawling archives provide the raw, messy context—the background noise of an era. The Librarian’s curated collections provide the focused, intelligible narratives within that noise. One is the landscape; the other is the carefully tended garden. The challenge for our time is not to choose a winner in this quiet contest, but to build bridges between these two philosophies, allowing the focused wisdom of the paste pot to find its necessary specimens within the spider’s vast, invaluable, and often impenetrable web.
Notes & further reading
A few pages I came back to while writing this:
- Gilbert, AZ
- The Glitch in the Family Photo: On Glitched Images as Public Records
- Peoria, AZ
- The Summer Fireflies and Vanishing Municipal Emails
- Surprise, AZ
- The Neutrality Mirage: Why 'Raw Data' is an Archive's Most Dangerous Fiction
- Elk Grove, CA
- Pasadena, CA
- New Haven, CT
- Stamford, CT
- Washington, DC
- one area's overview
- a practical rundown