The Neutrality Mirage: Why 'Raw Data' is an Archive's Most Dangerous Fiction

There’s a comforting story we tell ourselves about open data and digital preservation. It goes like this: if we can just capture the raw bits, store the pristine files, and hoard the unadulterated datasets, we have secured an objective record. The data, we imagine, is neutral ground. It’s the untouched bedrock upon which interpretations can be built and rebuilt. This belief is not just optimistic; it’s a dangerous fiction that obscures the fundamental subjectivity baked into every act of preservation.

The moment we decide something is worth preserving is the first editorial cut. A web crawler doesn’t wander the internet with impartial curiosity; it follows a script written by humans with priorities, biases, and blind spots. Which sites get archived daily, and which are captured once a decade? Which file formats are deemed ‘worthy’ of expensive, long-term conservation? The answers to these questions aren’t found in the data. They are found in the budgets, political pressures, and cultural assumptions of the institutions doing the archiving. The archive is not a reflection of the past; it is a curated argument about what mattered.

The Illusion of the Unmediated Capture

This becomes even more perilous when we treat things like public records datasets or massive social media dumps as ‘raw.’ A database of property records isn’t a simple mirror of reality. It is the product of centuries of law, bureaucratic categorization, and political boundaries. The fields in the spreadsheet—‘owner name,’ ‘land use,’ ‘value’—are not natural categories. They are administrative inventions that shape how we can even perceive the history of a place. To present this as raw is to hide its cooked nature.

Web archives magnify this. Capturing a webpage isn’t like taking a photograph of a static document. It’s a complex technical negotiation with a live system. The crawler’s view is one of billions of possible states, dependent on its location, the time of day, the user-agent string it pretends to be, and the scripts it can or cannot execute. The ‘raw data’ of a single archived page is already a specific, contingent performance of that page, frozen and mistaken for the whole. The interactive map that failed to load, the personalized ad that wasn’t served, the comment stream that was truncated—these aren’t accidents around the data. They *are* the data.

By clinging to the myth of neutrality, we risk creating a future that misunderstands our present. Researchers decades from now might query these ‘raw’ archives and believe they are seeing an unvarnished truth, unaware of the silent curators—the algorithms, the selection policies, the broken links deemed unimportant to fix—that shaped their view. The greatest threat to digital memory isn’t always bit rot; it’s the unexamined context rot that sets in when we forget that preservation is an argument, not a download.

The way forward isn’t to despair of bias, but to document it aggressively. True preservation means preserving the ‘why’ alongside the ‘what.’ It means saving the crawler’s configuration file, the data dictionary for the civic dataset, the minutes from the meeting where scope was debated. It means acknowledging that every archive is a lens, ground by its time. Only when we stop calling data ‘raw’ can we start to honestly taste the rich, complicated, and profoundly human stew we’ve actually saved.

Notes & further reading

A few pages I came back to while writing this: